
Senior Staff Infrastructure & Site Reliability Engineer Datacentre AI Engineering SA📣 Job Ad
| Contract Type | Full-time | |
| Workplace type | On-site | |
| Location | Riyadh |
Job Description
About the Opportunity at Qualcomm
Qualcomm is expanding its operations in Riyadh, Saudi Arabia, with significant investments in world-class computing and data center capabilities to support AI, cloud, and advanced connectivity initiatives aligned with Vision 2030. We are seeking a Senior Staff Infrastructure & Site Reliability Engineer to join our team in a full-time capacity. This role offers an opportunity to contribute to a growing technology hub, focusing on critical environments and the future of data center operations.
Role Overview and Impact
The Senior Staff Infrastructure & Site Reliability Engineer will be responsible for the design, operation, and continuous improvement of large-scale AI inference systems within a datacenter environment. This position is central to ensuring Qualcomm's AI infrastructure achieves high reliability, scalability, and readiness for advanced machine learning workloads. The role requires strong fundamentals in systems and software engineering, practical execution skills, and the ability to address complex problems independently while collaborating with hardware, software, and machine learning teams.
Key Responsibilities
- Design, deploy, and operate large-scale AI inference systems to support critical AI workloads.
- Ensure the reliability, availability, and scalability of Qualcomm's datacenter AI clusters.
- Develop and maintain software tools and support infrastructure for AI software stacks.
- Analyze software requirements and collaborate with architecture and hardware engineers to optimize AI workloads.
- Build, deploy, and operate components supporting LLM inference, agentic AI workflows, and AI services.
- Work with models, systems, and software teams to enhance model performance on AI100 deployments.
- Identify and implement optimizations for workloads running on multi-SoC and multi-card systems.
- Apply Site Reliability Engineering (SRE) fundamentals, including monitoring, alerting, incident response, and performance optimization.
- Support production ML systems using MLOps tools and operational best practices.
- Contribute to incident reviews, operational documentation, and continuous reliability improvements.
- Build and maintain observability tools, dashboards, and alerts for system health and reliability monitoring.
- Monitor infrastructure and services using tools such as Prometheus, Grafana, CloudWatch, and custom telemetry.
- Create and maintain technical documentation, runbooks, and knowledge-base articles.
- Develop automation to reduce manual operational tasks and improve system reliability.
- Support CI/CD pipelines for AI service and agent deployment.
- Apply Infrastructure-as-Code practices using tools such as Terraform and Ansible.
Required Qualifications and Experience
- Bachelor's degree in Engineering, Information Systems, Computer Science, or a related field, with 6+ years of Software Test Engineering or related work experience.
- Alternatively, a Master's degree in Engineering, Information Systems, Computer Science, or a related field, with 5+ years of Software Test Engineering or related work experience.
- Alternatively, a PhD in Engineering, Information Systems, Computer Science, or a related field, with 4+ years of Software Test Engineering or related work experience.
- A minimum of 2+ years of work experience in Software Test or System Test, including developing and automating test plans, and/or utilizing tools such as Source Code Control Systems, Continuous Integration Tools, and Bug Tracking Tools.
- While 10+ years of experience is generally expected for a Senior Staff role, references to a particular number of years are indicative. Candidates demonstrating equivalent experience and the ability to fulfill the principal duties and required competencies will be considered.
Technical Environment and Skills
The role involves working with a range of technologies and tools critical for AI infrastructure and SRE practices. This includes practical experience with:
- Large-scale AI inference systems, LLM inference, agentic AI workflows, and AI services.
- MLOps tools and operational best practices for production ML systems.
- Observability tools such as Prometheus, Grafana, CloudWatch, and custom telemetry.
- Infrastructure-as-Code practices using tools like Terraform and Ansible.
- Multi-SoC and multi-card systems for workload optimization.
Work Location and Type
This is a full-time position based in Riyadh, Saudi Arabia, supporting Qualcomm's expanding technology footprint in the region.
Requirements
- Requires +10 Years experience
Similar Jobs
You may also like
- Related Senior Staff Infrastructure & Site Reliability Engineer Datacentre AI Engineering SA Opportunities
- Sales Representative Jobs in Riyadh
- Customer Service Representative Jobs in Riyadh
- Seller Jobs in Riyadh
- Mechanical Engineer Jobs in Riyadh
- Electrical Engineer Jobs in Riyadh
- Other Job Fields in Riyadh
- Sales Representative Jobs in Riyadh
- Customer Service Representative Jobs in Riyadh
- Seller Jobs in Riyadh
- Mechanical Engineer Jobs in Riyadh
- Electrical Engineer Jobs in Riyadh
- Cashier Jobs in Riyadh
- Electronic Devices Technician Jobs in Riyadh
- Marketing Specialist Jobs in Riyadh
- IT Engineer Jobs in Riyadh
- Radiographer Jobs in Riyadh
- Explore Jobs Across Saudi Arabia
- Sales Specialist Jobs in Arar
- Process Control Specialist Jobs in Riyadh
- Project Management Specialist Jobs in Dammam
- Architect Jobs in Al-Kharj
- Intensivist Jobs in Najran
- Medical Optics Technician Jobs in Makkah
- Intermediate School Teacher of Islamic Studies Jobs in Ras Tannurah
- HVAC Technician Jobs in Arar
- Administrative Assistant Jobs in Makkah
- Sales Operations Specialist Jobs in Makkah
