img
Contract TypeFull-time
Workplace typeOn-site
LocationMakkah

Job Description

About the Role

S is seeking a Senior AI Backend Engineer - Agent Evaluation & Quality to join its team in Makkah, Makkah Province. This full-time role is central to ensuring the performance and reliability of production multi-agent systems that handle real work for a large user base.

Role Purpose and Impact

As these multi-agent systems grow, establishing confidence in their performance and preventing regressions is paramount. The successful candidate will be responsible for building core evaluation systems, including judges, test harnesses, and simulators, designed to accurately assess agent performance and identify failure points. This work will enable faster, more reliable agent deployments by fostering trust in the evaluation processes. A deep understanding of agent failure modes will also provide opportunities to contribute directly to agent development and improvement efforts.

Key Responsibilities

  • Own the evaluation stack: Design and build LLM-as-judge systems, calibrate them against human labels, and make agent quality measurable per-agent and per-failure-mode.
  • Implement a robust release gate: Build per-PR evaluation harnesses and regression detection wired into CI to enforce quality automatically, rather than through manual passes.
  • Develop user simulators to generate comprehensive test coverage and adversarial cases before they impact real users.
  • Translate production signals into improvements: Pipe real failures back into evaluation sets, allowing the system to compound and improve over time.
  • Partner with product teams to translate product requirements into concrete, measurable evaluation criteria.
  • Contribute to agent development: Grow into building and hardening the agents themselves, leveraging insights gained from their evaluation.

Required Qualifications and Experience

  • 5-10 years of software engineering experience, with recent hands-on LLM/agent work.
  • Strong software engineering fundamentals, including production experience with Python or Typescript (or similar), clean API and system design, testing, and CI/CD.
  • Hands-on LLM/agent experience: Demonstrated ability to build with LLMs, including agents, RAG, tool/function calling, and orchestration frameworks (such as LangGraph, LangChain, or equivalents), coupled with a deep understanding of their behavior and failure modes.
  • A measurement mindset: Adept at reasoning about metrics, calibration, and experiments, driven by a desire to quantify system effectiveness.
  • Production experience: Successfully managed and run LLM systems in production, addressing challenges related to reliability, latency, cost, and observability.

Desired Skills and Attributes

  • Direct experience evaluating LLM/agent systems, including offline/online evaluation, LLM-as-judge methodologies, and systematic regression testing.
  • Familiarity with observability tooling such as Arize, LangSmith, or similar platforms.
  • Arabic language or Natural Language Processing (NLP) experience.
  • E-commerce or merchant-facing product experience.

Application Information

Candidates meeting the above qualifications are encouraged to apply for this full-time position.


Requirements

  • Requires 5-10 Years experience

Similar Jobs