img
Contract TypeFull-time
Workplace typeOn-site
LocationRiyadh

Job Description

About the Opportunity at Qualcomm

Qualcomm is expanding its operations in Riyadh, Saudi Arabia, with significant investments in world-class computing and data center capabilities to support AI, cloud, and advanced connectivity initiatives aligned with Vision 2030. We are seeking a Senior Staff Infrastructure & Site Reliability Engineer to join our team in a full-time capacity. This role offers an opportunity to contribute to a growing technology hub, focusing on critical environments and the future of data center operations.

Role Overview and Impact

The Senior Staff Infrastructure & Site Reliability Engineer will be responsible for the design, operation, and continuous improvement of large-scale AI inference systems within a datacenter environment. This position is central to ensuring Qualcomm's AI infrastructure achieves high reliability, scalability, and readiness for advanced machine learning workloads. The role requires strong fundamentals in systems and software engineering, practical execution skills, and the ability to address complex problems independently while collaborating with hardware, software, and machine learning teams.

Key Responsibilities

  • Design, deploy, and operate large-scale AI inference systems to support critical AI workloads.
  • Ensure the reliability, availability, and scalability of Qualcomm's datacenter AI clusters.
  • Develop and maintain software tools and support infrastructure for AI software stacks.
  • Analyze software requirements and collaborate with architecture and hardware engineers to optimize AI workloads.
  • Build, deploy, and operate components supporting LLM inference, agentic AI workflows, and AI services.
  • Work with models, systems, and software teams to enhance model performance on AI100 deployments.
  • Identify and implement optimizations for workloads running on multi-SoC and multi-card systems.
  • Apply Site Reliability Engineering (SRE) fundamentals, including monitoring, alerting, incident response, and performance optimization.
  • Support production ML systems using MLOps tools and operational best practices.
  • Contribute to incident reviews, operational documentation, and continuous reliability improvements.
  • Build and maintain observability tools, dashboards, and alerts for system health and reliability monitoring.
  • Monitor infrastructure and services using tools such as Prometheus, Grafana, CloudWatch, and custom telemetry.
  • Create and maintain technical documentation, runbooks, and knowledge-base articles.
  • Develop automation to reduce manual operational tasks and improve system reliability.
  • Support CI/CD pipelines for AI service and agent deployment.
  • Apply Infrastructure-as-Code practices using tools such as Terraform and Ansible.

Required Qualifications and Experience

  • Bachelor's degree in Engineering, Information Systems, Computer Science, or a related field, with 6+ years of Software Test Engineering or related work experience.
  • Alternatively, a Master's degree in Engineering, Information Systems, Computer Science, or a related field, with 5+ years of Software Test Engineering or related work experience.
  • Alternatively, a PhD in Engineering, Information Systems, Computer Science, or a related field, with 4+ years of Software Test Engineering or related work experience.
  • A minimum of 2+ years of work experience in Software Test or System Test, including developing and automating test plans, and/or utilizing tools such as Source Code Control Systems, Continuous Integration Tools, and Bug Tracking Tools.
  • While 10+ years of experience is generally expected for a Senior Staff role, references to a particular number of years are indicative. Candidates demonstrating equivalent experience and the ability to fulfill the principal duties and required competencies will be considered.

Technical Environment and Skills

The role involves working with a range of technologies and tools critical for AI infrastructure and SRE practices. This includes practical experience with:

  • Large-scale AI inference systems, LLM inference, agentic AI workflows, and AI services.
  • MLOps tools and operational best practices for production ML systems.
  • Observability tools such as Prometheus, Grafana, CloudWatch, and custom telemetry.
  • Infrastructure-as-Code practices using tools like Terraform and Ansible.
  • Multi-SoC and multi-card systems for workload optimization.

Work Location and Type

This is a full-time position based in Riyadh, Saudi Arabia, supporting Qualcomm's expanding technology footprint in the region.


Requirements

  • Requires +10 Years experience

Similar Jobs