Agent Evaluation Engineer📣 Job Ad
| Contract Type | Part-time | |
| Workplace type | Remote | |
| Location | Saudi Arabia |
Job Description
About Mindrift and the Opportunity
Mindrift connects specialists with project-based AI opportunities for leading technology companies, focusing on the testing, evaluation, and improvement of AI systems. This is a part-time, freelance opportunity for a Freelance Agent Evaluation Engineer, distinct from permanent employment.
Role Purpose
The primary objective of this role is to contribute to building a dataset specifically designed for evaluating AI coding agents. This involves assessing how effectively an AI model handles real-world developer tasks. It is important to clarify that this position is not data labeling and not prompt engineering. While AI agents generate most of the code, the engineer's role is to guide and evaluate their solutions rather than writing code from scratch.
Key Responsibilities
- Build realistic developer environments: Create a virtual company complete with a codebase, infrastructure, and contextual information (such as tickets, documentation, and conversations) to establish a credible development history.
- Design tasks from intermediate states of these environments: Craft the initial prompt, define the criteria for a "solved" task, and ensure the task is solvable by an AI agent.
- Write tests that verify agent solutions: Develop tests that accurately accept all valid approaches while rejecting incorrect ones, ensuring they are neither excessively strict nor overly lenient.
- Iterate on tasks and tests based on Quality Assurance (QA) feedback: Review agent solutions, analyze instances of failure, and refine the evaluation process until it is fair and robust.
Required Expertise and Challenges
This role typically requires 5-10 years of relevant experience, given the complexity of the tasks. Frontier models are already proficient in coding, making the creation of tasks that genuinely challenge the best models non-trivial. A deep understanding of where models typically fail and what scenarios differentiate effective from ineffective solutions is necessary. Additionally, tasks often present multiple valid solutions, which makes writing tests that accurately accept all correct outcomes while rejecting incorrect ones a complex undertaking.
Work Arrangement and Compensation
This is a remote, part-time, freelance project designed to integrate flexibly with primary professional or academic commitments. Compensation is structured per accepted task, with earning potential up to $50 per hour. Participants will engage with advanced AI projects, contributing to their professional portfolio and influencing the development of future AI models.
Application Process
Interested candidates are invited to apply. The process involves passing qualification assessments, joining a project, completing assigned tasks, and subsequently receiving payment for accepted work.
Requirements
- Requires 5-10 Years experience
Similar Jobs
You may also like
- Related Agent Evaluation Engineer Opportunities
- Translator Jobs in Riyadh
- Barista Jobs in Riyadh
- Ride Operator Jobs in Riyadh
- Sweets Maker Jobs in Riyadh
- Chef Jobs in Riyadh
- Other Job Fields in
- Translator Jobs in Riyadh
- Barista Jobs in Riyadh
- Ride Operator Jobs in Riyadh
- Sweets Maker Jobs in Riyadh
- Chef Jobs in Riyadh
- Marketing Specialist Jobs in Riyadh
- Coffee Cashier Jobs in Riyadh
- HSE Officer Jobs in Riyadh
- Baker Jobs in Riyadh
- Communications Manager Jobs in Riyadh
- Explore Jobs Across Saudi Arabia
- Mechanical Engineering Technician Jobs in Riyadh
- HSE Engineer Jobs in Al-Kharj
- Receptionist Jobs in An Nuayriyah
- Dental Assistant Jobs in Dammam
- Seller Jobs in Buraydah
- Cleaning and Housekeeping Supervisor Jobs in Al-Ahsa
- Chemical Engineer Jobs in Jazan
- Solution Architect Jobs in Riyadh
- Tour Operator Jobs in Riyadh
- Restaurant Manager Jobs in Abha