img
Contract TypePart-time
Workplace typeRemote
LocationSaudi Arabia

Job Description

About Mindrift and the Opportunity

Mindrift connects specialists with project-based AI opportunities for leading technology companies, focusing on the testing, evaluation, and improvement of AI systems. This is a part-time, freelance opportunity for a Freelance Agent Evaluation Engineer, distinct from permanent employment.

Role Purpose

The primary objective of this role is to contribute to building a dataset specifically designed for evaluating AI coding agents. This involves assessing how effectively an AI model handles real-world developer tasks. It is important to clarify that this position is not data labeling and not prompt engineering. While AI agents generate most of the code, the engineer's role is to guide and evaluate their solutions rather than writing code from scratch.

Key Responsibilities

  • Build realistic developer environments: Create a virtual company complete with a codebase, infrastructure, and contextual information (such as tickets, documentation, and conversations) to establish a credible development history.
  • Design tasks from intermediate states of these environments: Craft the initial prompt, define the criteria for a "solved" task, and ensure the task is solvable by an AI agent.
  • Write tests that verify agent solutions: Develop tests that accurately accept all valid approaches while rejecting incorrect ones, ensuring they are neither excessively strict nor overly lenient.
  • Iterate on tasks and tests based on Quality Assurance (QA) feedback: Review agent solutions, analyze instances of failure, and refine the evaluation process until it is fair and robust.

Required Expertise and Challenges

This role typically requires 5-10 years of relevant experience, given the complexity of the tasks. Frontier models are already proficient in coding, making the creation of tasks that genuinely challenge the best models non-trivial. A deep understanding of where models typically fail and what scenarios differentiate effective from ineffective solutions is necessary. Additionally, tasks often present multiple valid solutions, which makes writing tests that accurately accept all correct outcomes while rejecting incorrect ones a complex undertaking.

Work Arrangement and Compensation

This is a remote, part-time, freelance project designed to integrate flexibly with primary professional or academic commitments. Compensation is structured per accepted task, with earning potential up to $50 per hour. Participants will engage with advanced AI projects, contributing to their professional portfolio and influencing the development of future AI models.

Application Process

Interested candidates are invited to apply. The process involves passing qualification assessments, joining a project, completing assigned tasks, and subsequently receiving payment for accepted work.


Requirements

  • Requires 5-10 Years experience

Similar Jobs