Agent Evaluation Engineer📣 إعلان
| نوع العقد | دوام جزئي | |
| طبيعة الوظيفة | عن بُعد | |
| الموقع | السعودية |
وصف الوظيفة
About the Role
Mindrift is seeking a Freelance Agent Evaluation Engineer to join projects focused on testing, evaluating, and improving AI systems for leading tech companies. This is a part-time, project-based freelance opportunity, not permanent employment. The role involves building datasets to evaluate AI coding agents and assess how well models handle real-world developer tasks. Compensation is paid per accepted task, with the potential to earn up to the equivalent of $50 per hour.
Key Responsibilities
- Build realistic developer environments, including a virtual company with a codebase, infrastructure, and context such as tickets, documentation, and conversations, to form a believable development history.
- Design tasks from intermediate states of these environments, crafting prompts, defining "solved" criteria, and ensuring tasks are solvable by an AI agent.
- Write tests to verify agent solutions, accepting all valid approaches and rejecting incorrect ones, maintaining an appropriate level of strictness.
- Iterate on tasks and tests based on QA feedback, reviewing agent solutions, analyzing failures, and refining evaluations for fairness and robustness.
Required Qualifications and Experience
- Minimum of 5 years of experience in software development.
- Minimum of 3 years of professional experience in related roles or domains, specifically in QA automation/testing or cybersecurity.
- Proficiency in Python (FastAPI), JavaScript/TypeScript (React), Docker, Postgres, Kafka, and Redis.
- Experience writing functional and integration tests.
- English proficiency at B2 level or higher.
Nature of the Work
This role is distinct from data labeling or prompt engineering. It does not involve writing code from scratch; instead, the focus is on guiding and evaluating AI agent-generated code. The work requires a deep understanding of where frontier AI models fail in coding tasks and the ability to design challenges that differentiate between effective and ineffective solutions. Developing tests that accurately accept all correct solutions while rejecting incorrect ones is a core aspect of this role.
Application Process
Candidates will apply, complete qualification assessments, join a project, complete assigned tasks, and receive payment upon task acceptance.
Benefits of this Freelance Opportunity
This freelance project offers the flexibility to work remotely and part-time, accommodating primary professional or academic commitments. Participants will engage with advanced AI projects, gaining experience that can enhance their professional portfolio and contribute to shaping future AI model understanding and communication within their field of expertise.
متطلبات الوظيفة
- تتطلب ٢-٥ سنوات خبرة
وظائف مشابهة
قد يعجبك أيضاً
- وظائف ذات صلة بـ Agent Evaluation Engineer
- وظائف محضر قهوة (باريستا) في الرياض
- وظائف أخصائي تطوير أعمال في الرياض
- وظائف موظف استقبال فندق في الرياض
- وظائف مهندس ذكاء اصطناعي في الرياض
- وظائف محاسب عام في الرياض
- مجالات وظيفية أخرى في
- وظائف محضر قهوة (باريستا) في الرياض
- وظائف أخصائي تطوير أعمال في الرياض
- وظائف موظف استقبال فندق في الرياض
- وظائف مهندس ذكاء اصطناعي في الرياض
- وظائف محاسب عام في الرياض
- وظائف Media Buyer في الرياض
- وظائف مندوب مبيعات في الرياض
- وظائف موظف استقبال في الرياض
- وظائف مندوب توصيل في الرياض
- وظائف مصور فوتوغرافي في الرياض
- استكشف الوظائف في أنحاء المملكة
- وظائف سائق شاحنة صغيرة في المدينة المنورة
- وظائف سائق سيارة خاص في خميس مشيط
- وظائف اخصائي علاج طبيعي في المدينة المنورة
- وظائف مراقب خدمات عامة في الأحساء
- وظائف Senior Accountant في الظهران
- وظائف معلم في المدينة المنورة
- وظائف مسؤول رعاية أطفال في الرياض
- وظائف طبيب أسنان عام في الدمام
- وظائف Videographer في الرياض
- وظائف مسؤول حملات سوشيال ميديا في جدة