Agent Evaluation Engineer
📣 إعلان| نوع العقد | دوام جزئي | |
| طبيعة الوظيفة | عن بُعد | |
| الموقع | السعودية |
وصف الوظيفة
About the Role
Mindrift is seeking a Freelance Agent Evaluation Engineer to join its team. This part-time, project-based role focuses on testing, evaluating, and improving AI systems for leading tech companies. The position is designed for specialists interested in contributing to advanced AI projects without permanent employment commitments.
Project Context
This role is integral to building a dataset for evaluating AI coding agents. The primary objective is to assess how effectively AI models handle real-world developer tasks. This is not a data labeling or prompt engineering role, nor does it involve writing code from scratch; instead, the focus is on guiding and evaluating AI-generated code.
Key Duties and Tasks
- Build realistic developer environments, including codebase, infrastructure, and contextual elements like tickets, documentation, and conversations, to form a believable development history.
- Design tasks from intermediate states of these environments, crafting prompts, defining "solved" criteria, and ensuring tasks are solvable by an AI agent.
- Write tests to verify agent solutions, accepting all valid approaches while rejecting incorrect ones, maintaining a balance that is neither too strict nor too lenient.
- Iterate on tasks and tests based on QA feedback, reviewing agent solutions, analyzing failures, and refining evaluations for fairness and robustness.
Challenges of the Role
This role presents unique challenges due to the advanced capabilities of frontier AI models in coding. Creating tasks that genuinely challenge these models requires a deep understanding of their failure points and scenarios that differentiate between effective and ineffective solutions. Additionally, designing tests that accommodate multiple valid solutions while correctly rejecting incorrect ones is a complex aspect of the evaluation process.
Qualifications and Experience
- 2-5 years of relevant experience.
- Ability to understand where AI models fail and identify scenarios that reveal differences between good and bad solutions.
- Proficiency in writing comprehensive tests that accept all correct solutions and reject incorrect ones.
Compensation and Work Arrangement
Compensation is paid per accepted task, with rates dependent on qualification tier and task completion efficiency, up to the equivalent of $50/hr. A faster pace of work can increase the effective hourly rate. This is a remote, part-time freelance opportunity designed to accommodate primary professional or academic commitments, offering valuable experience in advanced AI projects and the chance to influence future AI model development.
متطلبات الوظيفة
- تتطلب ٢-٥ سنوات خبرة
وظائف مشابهة
قد يعجبك أيضاً
- وظائف ذات صلة بـ Agent Evaluation Engineer
- وظائف Business Development Supervisor في الخبر
- وظائف بائع في الخبر
- وظائف موظف موارد بشرية في الخبر
- وظائف محضر قهوة (باريستا) في الخبر
- وظائف أخصائي تسويق في الخبر
- مجالات وظيفية أخرى في
- وظائف Business Development Supervisor في الخبر
- وظائف بائع في الخبر
- وظائف موظف موارد بشرية في الخبر
- وظائف محضر قهوة (باريستا) في الخبر
- وظائف أخصائي تسويق في الخبر
- وظائف Senior Accountant في الخبر
- وظائف محاسب زبائن (كاشير) في الخبر
- وظائف Demi Chef de Partie في الخبر
- وظائف صيدلي في الخبر
- وظائف مصمم جرافيك في الخبر
- استكشف الوظائف في أنحاء المملكة
- وظائف بائع هاتفي في الرياض
- وظائف Commercial Manager في الرياض
- وظائف سائق شاحنة صغيرة في تبوك
- وظائف أخصائي مبيعات في الرياض
- وظائف كهربائي مباني في تبوك
- وظائف مقيِّم عقاري في الرياض
- وظائف Massage Therapist في الخبر
- وظائف طبيب أسنان أطفال في مكة المكرمة
- وظائف مشغل محطة تحلية مياه في قبة
- وظائف أخصائي علاج طبيعي في نفي