img
نوع العقددوام جزئي
طبيعة الوظيفةعن بُعد
الموقعالسعودية

وصف الوظيفة

About the Role

Mindrift is seeking a Freelance Agent Evaluation Engineer to join projects focused on testing, evaluating, and improving AI systems for leading tech companies. This is a part-time, project-based freelance opportunity, not permanent employment. The role involves building datasets to evaluate AI coding agents and assess how well models handle real-world developer tasks. Compensation is paid per accepted task, with the potential to earn up to the equivalent of $50 per hour.

Key Responsibilities

  • Build realistic developer environments, including a virtual company with a codebase, infrastructure, and context such as tickets, documentation, and conversations, to form a believable development history.
  • Design tasks from intermediate states of these environments, crafting prompts, defining "solved" criteria, and ensuring tasks are solvable by an AI agent.
  • Write tests to verify agent solutions, accepting all valid approaches and rejecting incorrect ones, maintaining an appropriate level of strictness.
  • Iterate on tasks and tests based on QA feedback, reviewing agent solutions, analyzing failures, and refining evaluations for fairness and robustness.

Required Qualifications and Experience

  • Minimum of 5 years of experience in software development.
  • Minimum of 3 years of professional experience in related roles or domains, specifically in QA automation/testing or cybersecurity.
  • Proficiency in Python (FastAPI), JavaScript/TypeScript (React), Docker, Postgres, Kafka, and Redis.
  • Experience writing functional and integration tests.
  • English proficiency at B2 level or higher.

Nature of the Work

This role is distinct from data labeling or prompt engineering. It does not involve writing code from scratch; instead, the focus is on guiding and evaluating AI agent-generated code. The work requires a deep understanding of where frontier AI models fail in coding tasks and the ability to design challenges that differentiate between effective and ineffective solutions. Developing tests that accurately accept all correct solutions while rejecting incorrect ones is a core aspect of this role.

Application Process

Candidates will apply, complete qualification assessments, join a project, complete assigned tasks, and receive payment upon task acceptance.

Benefits of this Freelance Opportunity

This freelance project offers the flexibility to work remotely and part-time, accommodating primary professional or academic commitments. Participants will engage with advanced AI projects, gaining experience that can enhance their professional portfolio and contribute to shaping future AI model understanding and communication within their field of expertise.


متطلبات الوظيفة

  • تتطلب ٢-٥ سنوات خبرة

وظائف مشابهة