HPC Senior Systems Administrator📣 Job Ad
| Contract Type | Full-time | |
| Workplace type | On-site | |
| Location | Makkah |
Job Description
About the Role
KAUST (King Abdullah University of Science and Technology) is seeking a HPC Senior Systems Administrator to join its Supercomputing Laboratory (KSL) in Makkah, Saudi Arabia. This full-time role involves managing a large-scale HPC cluster, its storage systems, and networks, while providing critical support to researchers. The successful candidate will possess 5-10 years of relevant experience in HPC system administration.
Key Responsibilities
- Provide timely user support via multiple channels, maintaining high customer service standards.
- Install, configure, and manage HPC subsystems, including compute nodes, high-performance storage, InfiniBand, Ethernet, and configuration management tools (*, Ansible, Puppet).
- Deploy and manage cluster management software, monitoring tools, and supporting services for HPC clusters.
- Administer the Slurm workload manager, including QOS policies, accounts, accounting, and related automation scripts (Python, C++).
- Develop and maintain automation scripts in Bash and Python to streamline system administration tasks.
- Deploy and manage container environments (Singularity/Apptainer, Docker) for HPC workloads.
- Benchmark HPC system components (CPU, memory, InfiniBand, storage) periodically to ensure optimal performance and identify tuning opportunities.
- Enforce security best practices, including node hardening, kernel patching, and compliance across all systems.
- Manage parallel file systems such as Lustre, GPFS, Weka, or Vast, including performance tuning and capacity planning.
- Directly support research activities in computational science, engineering, data analysis, and AI/ML in collaboration with application support teams.
- Develop software tools and utilities as needed to support research projects on cluster systems.
- Drive proof-of-concept projects and technology evaluations, research industry best practices, and advocate for system enhancements.
- Coordinate with vendors and third-party service providers to report and resolve issues.
- Develop and maintain user documentation, standard operating procedures, and training materials.
Required Qualifications and Experience
- Demonstrated experience (5-10 years) troubleshooting complex hardware issues and documenting root cause analysis.
- Experience managing parallel storage systems (Lustre, GPFS, Weka, Vast, or similar).
- Proven experience benchmarking HPC system components (CPU, memory, InfiniBand, storage).
- Strong Linux system administration experience (RHEL, Rocky Linux, or CentOS) in large-scale HPC environments.
- Experience administering workload managers/schedulers (Slurm, LSF, or PBS).
- Proficiency with configuration management tools such as Ansible or Puppet.
- Familiarity with Kubernetes and container orchestration platforms is desirable.
Essential Skills and Competencies
- Expertise in supporting users of computational science, engineering, data analysis, and artificial intelligence applications and libraries in HPC environments.
- Proficiency with HPC applications and programming models (Fortran, C/C++, Python, MPI, OpenMP, CUDA, OpenACC).
- Demonstrated track record of managing complex HPC systems, including parallel file systems, job schedulers, InfiniBand/Ethernet networks, and monitoring systems.
- Knowledge of project management principles and practices.
- Strong analytical, problem-solving, and decision-making skills.
- Ability to proactively identify and implement system improvements, take initiative, and manage multiple concurrent projects to deliver high-quality results within deadlines.
- Excellent verbal and written communication skills in English, including the ability to prepare and deliver technical reports and presentations.
Work Environment and Collaboration
This role involves close collaboration with faculty, researchers, application support teams, and external partners. The administrator will coordinate with vendors and third-party service providers to resolve complex issues and drive them to closure. Effectiveness in multi-cultural, international work environments and proven cross-functional collaboration skills are essential.
Continuous Development
The successful candidate will be expected to stay at the forefront of HPC advancements through continuous learning, industry conferences, and professional collaboration. This includes driving benchmarking initiatives to inform future hardware procurement and advocating for system enhancements based on industry best practices.
Requirements
- Requires 5-10 Years experience
Similar Jobs
You may also like
- Related HPC Senior Systems Administrator Opportunities
- Florist Jobs in Riyadh
- Sales Representative Jobs in Riyadh
- Receptionist Jobs in Riyadh
- Hotel Housekeeping Supervisor Jobs in Riyadh
- Content Creator Jobs in Riyadh
- Other Job Fields in Makkah
- Sales Representative Jobs in Makkah
- Receptionist Jobs in Makkah
- Marketing Specialist Jobs in Makkah
- Business Development Specialist Jobs in Makkah
- Data Entry Agent Jobs in Makkah
- Cashier Jobs in Makkah
- Pastry Chef Jobs in Makkah
- Secretary Jobs in Makkah
- Photographer Jobs in Makkah
- Barista Jobs in Makkah
- Explore Jobs Across Saudi Arabia
- Mechanical Engineer Jobs in Al Baha
- Human Resources Specialist Jobs in Al Hafuf
- Graphic Designer Jobs in Al Khobar
- Cleaning worker Jobs in Abha
- Civil Engineer Jobs in Dhahran
- Dental Assistant Jobs in Huraymila
- Physiotherapy Specialist Jobs in Taif
- Environmental Engineer Jobs in Dhahran
- Nurse Specialist Jobs in Buraydah
- Medical Laboratory Technician Jobs in Riyadh