Agent Evaluation Engineer📣 إعلان
| نوع العقد | دوام جزئي | |
| طبيعة الوظيفة | عن بُعد | |
| الموقع | السعودية |
وصف الوظيفة
About Mindrift and the Opportunity
Mindrift connects specialists with project-based AI opportunities for leading technology companies, focusing on the testing, evaluation, and improvement of AI systems. This is a part-time, freelance opportunity for a Freelance Agent Evaluation Engineer, distinct from permanent employment.
Role Purpose
The primary objective of this role is to contribute to building a dataset specifically designed for evaluating AI coding agents. This involves assessing how effectively an AI model handles real-world developer tasks. It is important to clarify that this position is not data labeling and not prompt engineering. While AI agents generate most of the code, the engineer's role is to guide and evaluate their solutions rather than writing code from scratch.
Key Responsibilities
- Build realistic developer environments: Create a virtual company complete with a codebase, infrastructure, and contextual information (such as tickets, documentation, and conversations) to establish a credible development history.
- Design tasks from intermediate states of these environments: Craft the initial prompt, define the criteria for a "solved" task, and ensure the task is solvable by an AI agent.
- Write tests that verify agent solutions: Develop tests that accurately accept all valid approaches while rejecting incorrect ones, ensuring they are neither excessively strict nor overly lenient.
- Iterate on tasks and tests based on Quality Assurance (QA) feedback: Review agent solutions, analyze instances of failure, and refine the evaluation process until it is fair and robust.
Required Expertise and Challenges
This role typically requires 5-10 years of relevant experience, given the complexity of the tasks. Frontier models are already proficient in coding, making the creation of tasks that genuinely challenge the best models non-trivial. A deep understanding of where models typically fail and what scenarios differentiate effective from ineffective solutions is necessary. Additionally, tasks often present multiple valid solutions, which makes writing tests that accurately accept all correct outcomes while rejecting incorrect ones a complex undertaking.
Work Arrangement and Compensation
This is a remote, part-time, freelance project designed to integrate flexibly with primary professional or academic commitments. Compensation is structured per accepted task, with earning potential up to $50 per hour. Participants will engage with advanced AI projects, contributing to their professional portfolio and influencing the development of future AI models.
Application Process
Interested candidates are invited to apply. The process involves passing qualification assessments, joining a project, completing assigned tasks, and subsequently receiving payment for accepted work.
متطلبات الوظيفة
- تتطلب ٥-١٠ سنوات خبرة
وظائف مشابهة
قد يعجبك أيضاً
- وظائف ذات صلة بـ Agent Evaluation Engineer
- وظائف شيف قسم المخبوزات في الرياض
- وظائف painter في الرياض
- وظائف General Accountant في الرياض
- وظائف مندوب مبيعات في الرياض
- وظائف بائع في الرياض
- مجالات وظيفية أخرى في
- وظائف شيف قسم المخبوزات في الرياض
- وظائف painter في الرياض
- وظائف General Accountant في الرياض
- وظائف مندوب مبيعات في الرياض
- وظائف بائع في الرياض
- وظائف مهندس مبيعات تقنية في الرياض
- وظائف أخصائي تطوير أعمال في الرياض
- وظائف أخصائي تسويق في الرياض
- وظائف تنفيذي مبيعات في الرياض
- وظائف Content Creator في الرياض
- استكشف الوظائف في أنحاء المملكة
- وظائف مراقب اجتماعي في الرس
- وظائف أخصائي اضطرابات تخاطب في جازان
- وظائف منقذ سباحة في جازان
- وظائف سائق شاحنة صغيرة في المدينة المنورة
- وظائف فني كهربائي أنظمة حماية كهربائية في نفي
- وظائف Butcher في املج
- وظائف مدرب معتمد في الدمام
- وظائف Risk Manager في تبوك
- وظائف مهندس زراعي في ينبع
- وظائف أمين صندوق في جدة