img
Contract TypeFull-time
Workplace typeOn-site
LocationRiyadh

Job Description

About the Role

Think is seeking an ML Systems Software Engineer to join their team in Riyadh. This full-time role focuses on the critical interface between machine learning models and hardware, where system-level decisions directly impact computational efficiency. The position requires 2-5 years of relevant experience.

Role Context and Impact

This role operates at the core of ILM, focusing on the layer where scheduling, memory management, and kernel path choices determine whether silicon is actively computing or waiting. The engineer will be instrumental in optimising the performance of ML systems by addressing bottlenecks and ensuring efficient accelerator utilisation.

Key Responsibilities

  • Build and optimise the serving path, including batching, memory management, cache behaviour, and scheduling under real concurrency.
  • Profile end-to-end system performance to identify actual bottlenecks.
  • Enable efficient sharing of accelerators across various silicon capacities and generations.
  • Work closely with the runtime environment, including drivers, kernels, memory allocators, and telemetry systems.
  • Develop measurement harnesses to validate performance on real hardware.
  • Drive changes from initial hypothesis through benchmarking to production deployment.

Required Qualifications and Skills

  • Four or more years of experience in systems or ML infrastructure engineering.
  • Strong proficiency in Python and a systems language such as C++ or Rust.
  • Demonstrated experience optimising inference or training throughput, with the ability to articulate the source of performance gains.
  • Understanding of accelerator memory hierarchies, kernel launch behaviour, and performance profiling.
  • Comfort working directly on bare metal systems, outside of managed service environments.

Beneficial Skills and Experience

  • Experience with CUDA, ROCm, Triton, or comparable kernel-level development.
  • Contributions to projects such as vLLM, TensorRT-LLM, or SGLang.
  • Experience with distributed serving and multi-accelerator sharding.
  • Published benchmarks or systems-related work.

Work Environment

This is a full-time position based in Riyadh, focusing on advanced ML systems engineering within a dynamic technical environment.


Requirements

  • Requires 5-10 Years experience

Similar Jobs