img
نوع العقددوام كامل
طبيعة الوظيفةبالموقع
الموقعالرياض

وصف الوظيفة

About the Role

TAWANTECH is seeking an Expert Site Reliability Engineer to join their team in Riyadh, Saudi Arabia. This full-time role focuses on enhancing the reliability, availability, scalability, and operational resilience of critical technology services through advanced software engineering and reliability practices. The position requires 5-10 years of experience in relevant fields.

Purpose of the Position

The primary purpose of this role is to drive the operational resilience of critical technology services. This involves applying advanced software engineering, automation, observability, and reliability engineering practices to ensure high availability and performance across all systems.

Key Responsibilities

  • Define and implement advanced reliability engineering practices for critical technology services.
  • Establish and monitor Service Level Indicators (SLIs), Service Level Objectives (SLOs), and reliability targets.
  • Design automation solutions to reduce manual operational activities and improve system resilience.
  • Develop and enhance monitoring, observability, alerting, and incident detection capabilities.
  • Lead technical analysis and resolution of complex production incidents.
  • Conduct root-cause analysis and drive permanent corrective and preventive actions.
  • Design solutions to improve system availability, scalability, capacity, and disaster resilience.
  • Identify reliability risks and recommend architectural and engineering improvements.
  • Drive performance engineering and capacity planning for critical services.
  • Provide advanced technical guidance and mentorship on SRE practices.
  • Promote automation and engineering approaches to reduce operational toil and improve service reliability.

Qualifications and Experience

  • Bachelor's degree in Computer Science, Software Engineering, IT, or a related field.
  • 5+ years of experience in Site Reliability Engineering, DevOps, Platform Engineering, or similar roles.
  • Strong experience with cloud platforms, Kubernetes, and production environments.
  • Strong knowledge of monitoring, observability, alerting, SLIs, SLOs, and reliability metrics.
  • Hands-on experience with automation, scripting, CI/CD, and Infrastructure as Code (Terraform).
  • Proven experience in complex incident management, troubleshooting, and Root Cause Analysis (RCA).
  • Strong understanding of high availability, scalability, performance engineering, capacity planning, and disaster recovery.
  • Experience driving reliability improvements and reducing operational toil through automation.
  • Strong analytical, problem-solving, and technical leadership skills.
  • Experience in Banking, FinTech, or Payment environments is preferred.

Work Environment

This is a full-time position based in Riyadh, Saudi Arabia, offering an opportunity to contribute to critical technology services within a professional setting.


متطلبات الوظيفة

  • تتطلب ٥-١٠ سنوات خبرة

وظائف مشابهة