img
Contract TypeFull-time
Workplace typeOn-site
LocationRiyadh

Job Description

About the Role

Mirai is seeking a Senior Data Engineer to join its team in Riyadh. This full-time role is central to the development of Generative AI products, focusing on the end-to-end data layer. The successful candidate will be responsible for the pipelines, transformations, and delivery of data to retrieval systems, agents, and analytics platforms.

Role Context and Objectives

The quality of Generative AI products is directly dependent on the underlying data. This position involves owning the entire data lifecycle, from ingestion to consumption, with the objective of establishing a single, governed data source that every consumer can rely on. The work operates within an AWS environment, ensuring data reliability and accessibility. The team is small and spans several languages, requiring individuals to manage their pipelines and contribute to setting data layer standards.

Key Responsibilities

  • Build and operate batch and streaming pipelines that move data from source systems into the data lake and through to the warehouse, owning the layers in between from raw to curated, along with their schema, quality, and lineage.
  • Develop the data layer behind retrieval systems, including source connectors, document parsing, chunking, embedding generation, and vector indexing, with provisions for re-embedding when content changes.
  • Model curated, query-ready datasets and metrics to ensure AI and analytics consumers work from one definition.
  • Implement quality checks, validation, and monitoring to identify data problems before they impact models or users.
  • Apply access controls, including row and column level rules, PII handling, and entitlement-aware datasets, enforced as close to query time as the stack allows.
  • Collaborate with platform and DevOps engineers to expose data and retrieval as documented, dependable services.
  • Manage storage, compute, and query costs, with particular attention to the cost of embedding and vector workloads.
  • Participate in code reviews, write documentation, and help shape the team's approach to building its data layer.

Required Qualifications and Experience

  • A minimum of eight or more years of experience in data engineering.
  • Demonstrated hands-on experience building data for AI or Machine Learning systems, such as retrieval, embeddings, or feature data.
  • Proven experience in preparing data for Large Language Models (LLMs) or agents, including work around chunking, embeddings, indexing, and keeping content current.

Technical Environment

The data infrastructure operates on AWS. Candidates should be familiar with cloud-native data services and best practices for managing data at scale within this environment. A strong understanding of data preparation techniques for AI systems, beyond traditional reporting, is essential for this role.

Work Type and Location

This is a full-time position based in Riyadh. The role requires an individual who can operate effectively within a small, collaborative team environment.


Requirements

  • Requires 5-10 Years experience

Similar Jobs