
Senior Staff Infrastructure & Site Reliability Engineer Datacentre AI Engineering SA📣 إعلان
| نوع العقد | دوام كامل | |
| طبيعة الوظيفة | بالموقع | |
| الموقع | الرياض |
وصف الوظيفة
About the Opportunity at Qualcomm
Qualcomm is expanding its operations in Riyadh, Saudi Arabia, with significant investments in world-class computing and data center capabilities to support AI, cloud, and advanced connectivity initiatives aligned with Vision 2030. We are seeking a Senior Staff Infrastructure & Site Reliability Engineer to join our team in a full-time capacity. This role offers an opportunity to contribute to a growing technology hub, focusing on critical environments and the future of data center operations.
Role Overview and Impact
The Senior Staff Infrastructure & Site Reliability Engineer will be responsible for the design, operation, and continuous improvement of large-scale AI inference systems within a datacenter environment. This position is central to ensuring Qualcomm's AI infrastructure achieves high reliability, scalability, and readiness for advanced machine learning workloads. The role requires strong fundamentals in systems and software engineering, practical execution skills, and the ability to address complex problems independently while collaborating with hardware, software, and machine learning teams.
Key Responsibilities
- Design, deploy, and operate large-scale AI inference systems to support critical AI workloads.
- Ensure the reliability, availability, and scalability of Qualcomm's datacenter AI clusters.
- Develop and maintain software tools and support infrastructure for AI software stacks.
- Analyze software requirements and collaborate with architecture and hardware engineers to optimize AI workloads.
- Build, deploy, and operate components supporting LLM inference, agentic AI workflows, and AI services.
- Work with models, systems, and software teams to enhance model performance on AI100 deployments.
- Identify and implement optimizations for workloads running on multi-SoC and multi-card systems.
- Apply Site Reliability Engineering (SRE) fundamentals, including monitoring, alerting, incident response, and performance optimization.
- Support production ML systems using MLOps tools and operational best practices.
- Contribute to incident reviews, operational documentation, and continuous reliability improvements.
- Build and maintain observability tools, dashboards, and alerts for system health and reliability monitoring.
- Monitor infrastructure and services using tools such as Prometheus, Grafana, CloudWatch, and custom telemetry.
- Create and maintain technical documentation, runbooks, and knowledge-base articles.
- Develop automation to reduce manual operational tasks and improve system reliability.
- Support CI/CD pipelines for AI service and agent deployment.
- Apply Infrastructure-as-Code practices using tools such as Terraform and Ansible.
Required Qualifications and Experience
- Bachelor's degree in Engineering, Information Systems, Computer Science, or a related field, with 6+ years of Software Test Engineering or related work experience.
- Alternatively, a Master's degree in Engineering, Information Systems, Computer Science, or a related field, with 5+ years of Software Test Engineering or related work experience.
- Alternatively, a PhD in Engineering, Information Systems, Computer Science, or a related field, with 4+ years of Software Test Engineering or related work experience.
- A minimum of 2+ years of work experience in Software Test or System Test, including developing and automating test plans, and/or utilizing tools such as Source Code Control Systems, Continuous Integration Tools, and Bug Tracking Tools.
- While 10+ years of experience is generally expected for a Senior Staff role, references to a particular number of years are indicative. Candidates demonstrating equivalent experience and the ability to fulfill the principal duties and required competencies will be considered.
Technical Environment and Skills
The role involves working with a range of technologies and tools critical for AI infrastructure and SRE practices. This includes practical experience with:
- Large-scale AI inference systems, LLM inference, agentic AI workflows, and AI services.
- MLOps tools and operational best practices for production ML systems.
- Observability tools such as Prometheus, Grafana, CloudWatch, and custom telemetry.
- Infrastructure-as-Code practices using tools like Terraform and Ansible.
- Multi-SoC and multi-card systems for workload optimization.
Work Location and Type
This is a full-time position based in Riyadh, Saudi Arabia, supporting Qualcomm's expanding technology footprint in the region.
متطلبات الوظيفة
- تتطلب اكثر من ١٠ سنوات خبرة
وظائف مشابهة
قد يعجبك أيضاً
- وظائف ذات صلة بـ Senior Staff Infrastructure & Site Reliability Engineer Datacentre AI Engineering SA
- وظائف مندوب مبيعات في الرياض
- وظائف ممثل خدمة عملاء في الرياض
- وظائف بائع في الرياض
- وظائف Mechanical Engineer في الرياض
- وظائف Electrical Engineer في الرياض
- مجالات وظيفية أخرى في الرياض
- وظائف مندوب مبيعات في الرياض
- وظائف ممثل خدمة عملاء في الرياض
- وظائف بائع في الرياض
- وظائف Mechanical Engineer في الرياض
- وظائف Electrical Engineer في الرياض
- وظائف Cashier في الرياض
- وظائف فني أجهزة إلكترونية في الرياض
- وظائف أخصائي تسويق في الرياض
- وظائف مهندس تقنية معلومات في الرياض
- وظائف فني أشعة في الرياض
- استكشف الوظائف في أنحاء المملكة
- وظائف مهندس الصحة والسلامة والبيئة في الظهران
- وظائف فني ميكانيكي تمديدات صحية وتدفئة وغاز في الدمام
- وظائف موظف استقبال في الجبيل
- وظائف مهندس أجهزة طبية في مكة المكرمة
- وظائف مهندس زراعي في جدة
- وظائف فني تدفئة وتهوية وتكييف في العلا
- وظائف مشغل آلة خياطة في الرياض
- وظائف IT Specialist في جدة
- وظائف مراقب حركة مركبات في جدة
- وظائف سائق سيارة خاص في حائل
