Agent Evaluation Engineer📣 إعلان
| نوع العقد | دوام جزئي | |
| طبيعة الوظيفة | عن بُعد | |
| الموقع | السعودية |
وصف الوظيفة
About Mindrift and the Opportunity
Mindrift connects specialists with project-based AI opportunities for leading technology companies, focusing on the testing, evaluation, and improvement of AI systems. This is a part-time, freelance opportunity for a Freelance Agent Evaluation Engineer, distinct from permanent employment.
Role Purpose
The primary objective of this role is to contribute to building a dataset specifically designed for evaluating AI coding agents. This involves assessing how effectively an AI model handles real-world developer tasks. It is important to clarify that this position is not data labeling and not prompt engineering. While AI agents generate most of the code, the engineer's role is to guide and evaluate their solutions rather than writing code from scratch.
Key Responsibilities
- Build realistic developer environments: Create a virtual company complete with a codebase, infrastructure, and contextual information (such as tickets, documentation, and conversations) to establish a credible development history.
- Design tasks from intermediate states of these environments: Craft the initial prompt, define the criteria for a "solved" task, and ensure the task is solvable by an AI agent.
- Write tests that verify agent solutions: Develop tests that accurately accept all valid approaches while rejecting incorrect ones, ensuring they are neither excessively strict nor overly lenient.
- Iterate on tasks and tests based on Quality Assurance (QA) feedback: Review agent solutions, analyze instances of failure, and refine the evaluation process until it is fair and robust.
Required Expertise and Challenges
This role typically requires 5-10 years of relevant experience, given the complexity of the tasks. Frontier models are already proficient in coding, making the creation of tasks that genuinely challenge the best models non-trivial. A deep understanding of where models typically fail and what scenarios differentiate effective from ineffective solutions is necessary. Additionally, tasks often present multiple valid solutions, which makes writing tests that accurately accept all correct outcomes while rejecting incorrect ones a complex undertaking.
Work Arrangement and Compensation
This is a remote, part-time, freelance project designed to integrate flexibly with primary professional or academic commitments. Compensation is structured per accepted task, with earning potential up to $50 per hour. Participants will engage with advanced AI projects, contributing to their professional portfolio and influencing the development of future AI models.
Application Process
Interested candidates are invited to apply. The process involves passing qualification assessments, joining a project, completing assigned tasks, and subsequently receiving payment for accepted work.
متطلبات الوظيفة
- تتطلب ٥-١٠ سنوات خبرة
وظائف مشابهة
قد يعجبك أيضاً
- وظائف ذات صلة بـ Agent Evaluation Engineer
- وظائف Translator في الرياض
- وظائف محضر قهوة (باريستا) في الرياض
- وظائف مشغل ألعاب في الرياض
- وظائف شيف حلويات في الرياض
- وظائف طاهي (شيف) في الرياض
- مجالات وظيفية أخرى في
- وظائف Translator في الرياض
- وظائف محضر قهوة (باريستا) في الرياض
- وظائف مشغل ألعاب في الرياض
- وظائف شيف حلويات في الرياض
- وظائف طاهي (شيف) في الرياض
- وظائف أخصائي تسويق في الرياض
- وظائف كاشير قهوة في الرياض
- وظائف HSE Officer في الرياض
- وظائف Baker في الرياض
- وظائف Communications Manager في الرياض
- استكشف الوظائف في أنحاء المملكة
- وظائف Organizational Development Specialist في جدة
- وظائف أخصائي تمريض في المبرز
- وظائف أخصائي استمرارية الأعمال في الرياض
- وظائف Quality Manager في الخبر
- وظائف فني كهروميكانيك في جازان
- وظائف Internal Auditor في الدمام
- وظائف فني أشعة في جدة
- وظائف منقذ سباحة في جازان
- وظائف أخصائي قبول وتسجيل في الدمام
- وظائف فني مساحة في الخبر