Agent Evaluation Engineer
📣 Job Ad| Contract Type | Part-time | |
| Workplace type | Remote | |
| Location | Saudi Arabia |
Job Description
About the Role
Mindrift is seeking a Freelance Agent Evaluation Engineer to join its team. This part-time, project-based role focuses on testing, evaluating, and improving AI systems for leading tech companies. The position is designed for specialists interested in contributing to advanced AI projects without permanent employment commitments.
Project Context
This role is integral to building a dataset for evaluating AI coding agents. The primary objective is to assess how effectively AI models handle real-world developer tasks. This is not a data labeling or prompt engineering role, nor does it involve writing code from scratch; instead, the focus is on guiding and evaluating AI-generated code.
Key Duties and Tasks
- Build realistic developer environments, including codebase, infrastructure, and contextual elements like tickets, documentation, and conversations, to form a believable development history.
- Design tasks from intermediate states of these environments, crafting prompts, defining "solved" criteria, and ensuring tasks are solvable by an AI agent.
- Write tests to verify agent solutions, accepting all valid approaches while rejecting incorrect ones, maintaining a balance that is neither too strict nor too lenient.
- Iterate on tasks and tests based on QA feedback, reviewing agent solutions, analyzing failures, and refining evaluations for fairness and robustness.
Challenges of the Role
This role presents unique challenges due to the advanced capabilities of frontier AI models in coding. Creating tasks that genuinely challenge these models requires a deep understanding of their failure points and scenarios that differentiate between effective and ineffective solutions. Additionally, designing tests that accommodate multiple valid solutions while correctly rejecting incorrect ones is a complex aspect of the evaluation process.
Qualifications and Experience
- 2-5 years of relevant experience.
- Ability to understand where AI models fail and identify scenarios that reveal differences between good and bad solutions.
- Proficiency in writing comprehensive tests that accept all correct solutions and reject incorrect ones.
Compensation and Work Arrangement
Compensation is paid per accepted task, with rates dependent on qualification tier and task completion efficiency, up to the equivalent of $50/hr. A faster pace of work can increase the effective hourly rate. This is a remote, part-time freelance opportunity designed to accommodate primary professional or academic commitments, offering valuable experience in advanced AI projects and the chance to influence future AI model development.
Requirements
- Requires 2-5 Years experience
Similar Jobs
You may also like
- Related Agent Evaluation Engineer Opportunities
- Business Development Supervisor Jobs in Al Khobar
- Seller Jobs in Al Khobar
- Human Resources Clerk Jobs in Al Khobar
- Barista Jobs in Al Khobar
- Marketing Specialist Jobs in Al Khobar
- Other Job Fields in
- Business Development Supervisor Jobs in Al Khobar
- Seller Jobs in Al Khobar
- Human Resources Clerk Jobs in Al Khobar
- Barista Jobs in Al Khobar
- Marketing Specialist Jobs in Al Khobar
- Sales Representative Jobs in Al Khobar
- Senior Accountant Jobs in Al Khobar
- Cashier Jobs in Al Khobar
- Demi Chef de Partie Jobs in Al Khobar
- Pharmacist Jobs in Al Khobar
- Explore Jobs Across Saudi Arabia
- Sales Representative Jobs in Al Qatif
- Service Management Specialist Jobs in Riyadh
- Call Center Agent Jobs in Al Khobar
- Mechanical Engineer Jobs in Khamis Mushayt
- Sales Representative Jobs in Jazan
- Technical Supervisor Jobs in Al Majmaah
- Mechanical Engineer Jobs in Nifi
- Cost Controller Jobs in Riyadh
- Chemist Jobs in Khamis Mushayt
- Psychological Therapist Jobs in Bishah