img
الراتب15,000SR / شهرياً
نوع العقدعمل مؤقت
طبيعة الوظيفةبالموقع
الموقعالرياض

وصف الوظيفة

About the Role

Throne Solutions is seeking a Site Reliability Engineer (SRE) to join our technical team in Riyadh, Saudi Arabia. This is a full-time freelance contract position, offering a salary of SAR 15,000 – 17,000 per month. The role is suitable for candidates with 0-1 years of experience in a relevant field. The SRE will be responsible for ensuring the availability, reliability, scalability, performance, and security of enterprise applications and infrastructure. This role combines systems engineering, cloud operations, automation, monitoring, incident management, and DevOps practices.

Key Responsibilities

The Site Reliability Engineer will perform a range of duties, including:

  • Ensure high availability, reliability, and performance of production systems, including monitoring infrastructure and application health.
  • Identify and resolve performance, availability, and capacity issues, participating in incident, problem, and change management processes.
  • Provide L2/L3 support for critical production incidents, perform root-cause analysis (RCA), and implement permanent corrective actions.
  • Implement and maintain comprehensive monitoring and alerting solutions, developing dashboards and meaningful alerts for proactive issue detection.
  • Administer and support cloud infrastructure across AWS, Microsoft Azure, or Google Cloud, managing compute, storage, networking, IAM, and cloud monitoring services.
  • Deploy, manage, and troubleshoot containerized applications and Kubernetes clusters in production environments.
  • Automate repetitive operational and infrastructure tasks using scripting (Python, Bash, PowerShell) and Infrastructure as Code tools (Terraform, Ansible).
  • Support and maintain CI/CD pipelines using platforms like Jenkins, GitLab CI/CD, GitHub Actions, or Azure DevOps, automating application and infrastructure deployment.
  • Respond to production incidents within defined SLAs, participate in troubleshooting high-priority incidents, and prepare incident reports and post-incident reviews.

Technical Expertise and Tools

Candidates should possess hands-on experience with the following technologies and domains:

  • Operating Systems: Linux/Windows systems.
  • Cloud Platforms: AWS, Microsoft Azure, or Google Cloud.
  • Containerization: Kubernetes and Docker.
  • CI/CD: Jenkins, GitLab CI/CD, GitHub Actions, Azure DevOps, or similar platforms.
  • Monitoring & Observability: Prometheus, Grafana, ELK/Elastic Stack, Splunk, Datadog, Nagios, or Zabbix.
  • Automation & IaC: Python, Bash, PowerShell, Terraform, and Ansible.

Required Experience

The ideal candidate will demonstrate experience in:

  • Supporting business-critical production environments.
  • Hands-on experience with cloud infrastructure and automation.
  • Troubleshooting complex production incidents.
  • Monitoring, alerting, logging, and observability practices.
  • Practical experience with Kubernetes and containerized environments.
  • Working with CI/CD pipelines and Infrastructure as Code.
  • A strong understanding of reliability, scalability, and high-availability concepts.

Candidate Profile

We are looking for a candidate who exhibits:

  • Strong troubleshooting and analytical skills.
  • A proactive approach to reliability and automation.
  • Ability to work effectively during high-severity production incidents.
  • Strong communication and documentation skills.
  • Ability to collaborate with development, network, security, database, and infrastructure teams.
  • Ability to work independently in an onsite customer environment.
  • Willingness to participate in scheduled maintenance and on-call support when required.

متطلبات الوظيفة

  • لا تتطلب خبرة

وظائف مشابهة