img
نوع العقددوام كامل
طبيعة الوظيفةبالموقع
الموقعالأحساء

وصف الوظيفة

About TestCrew and the Role

TestCrew | Quality Engineering & Software Testing is seeking a Site Reliability Engineer (SRE) for an upcoming enterprise engagement in Al-Ahsa, located in the Eastern region. This full-time position is open to candidates with 2-5 years of relevant experience, focusing on ensuring the stability and performance of critical IT systems.

Role Overview

The Site Reliability Engineer will undertake a hands-on operational role, responsible for the availability, reliability, performance, and security of critical IT systems. This involves proactive monitoring, automation, and disciplined incident response. The successful candidate will collaborate closely with development, operations, and security teams to maintain highly available services, streamline operations through automation, and drive continuous service improvement.

Key Responsibilities

  • Monitor and manage enterprise infrastructure, applications, and services to ensure high availability, stability, and optimal performance.
  • Operate and maintain monitoring and observability platforms to detect incidents, analyze trends, and respond proactively to system issues.
  • Design and implement monitoring strategies, dashboards, alerts, and performance metrics to improve operational visibility.
  • Perform root cause analysis (RCA) for incidents and implement preventive measures to reduce recurring issues.
  • Automate operational tasks, deployments, and maintenance activities using scripting and infrastructure automation tools.
  • Manage CI/CD pipelines and support release management processes to enable reliable and efficient software deployments.
  • Optimize system performance while ensuring compliance with service level agreements (SLAs), operational standards, and best practices.
  • Support business continuity, disaster recovery, and operational resilience initiatives.
  • Implement security best practices, including system hardening, patch management, and secure operational procedures.
  • Collaborate with development, infrastructure, security, and IT operations teams to resolve complex technical issues.
  • Maintain operational documentation, runbooks, knowledge articles, and incident reports.
  • Participate in major incident management, post-incident reviews, and continuous improvement initiatives.

Candidate Profile and Requirements

  • 2-5 years of experience in a Site Reliability Engineering or a similar operational role.
  • Strong ownership mindset with a proactive, reliability-first approach.
  • Excellent analytical and problem-solving skills.
  • Ability to remain calm and methodical during critical incidents.
  • Strong communication and collaboration skills.
  • Ability to work effectively in a fast-paced enterprise environment.

Work Environment and Conditions

This role requires candidates to work onsite at the client location in Al-Ahsa. Participation in on-call support and shift rotations is also a requirement for this position.

Application Information

This full-time position offers an opportunity to contribute to critical system reliability within an enterprise environment. Salary details will be discussed during the interview process.


متطلبات الوظيفة

  • لا تتطلب خبرة

وظائف مشابهة