img
نوع العقددوام كامل
طبيعة الوظيفةبالموقع
الموقعالسعودية

وصف الوظيفة

About the Role

King Abdullah University of Science and Technology (KAUST) is seeking a Senior HPC Systems Administrator to join the KAUST Supercomputing Laboratory (KSL). This full-time role involves managing an HPC cluster of approximately 600 CPU and GPU nodes, HPC storage systems, InfiniBand and Ethernet networks, and addressing day-to-day operational issues. The position provides broad support to researchers and end-users across various computational domains.

Key Responsibilities

  • Provide timely user support via telephone, walk-in, email, and ticketing system, maintaining high customer service standards.
  • Install, configure, and manage HPC subsystems, including compute nodes, high-performance storage systems, InfiniBand, Ethernet, and configuration management tools (*, Ansible, Puppet).
  • Deploy and manage cluster management software, monitoring tools, and supporting services for operating HPC clusters.
  • Install and administer the Slurm workload manager, managing QOS policies, accounts, accounting, and related automation scripts (Python and C++).
  • Develop and maintain automation scripts in Bash and Python to streamline system administration tasks.
  • Deploy and manage container environments (Singularity/Apptainer, Docker) for HPC workloads.
  • Benchmark HPC system components (CPU, memory, InfiniBand, storage) periodically to ensure optimal performance and identify tuning opportunities.
  • Enforce security best practices, including node hardening, kernel patching, and compliance across all systems.
  • Manage parallel file systems such as Lustre, GPFS, Weka, or Vast, including performance tuning and capacity planning.
  • Directly support research activities in computational science, engineering, data analysis, and AI/ML in collaboration with faculty, researchers, and partners.
  • Develop software tools and utilities as needed to support research projects on cluster systems and subsystems.
  • Drive proof-of-concept projects and technology evaluations, research industry best practices, and advocate system enhancements.
  • Coordinate with vendors and third-party service providers to report and resolve issues.
  • Develop and maintain user documentation, standard operating procedures, and training materials in the internal wiki.
  • Stay informed of HPC advancements through continuous learning, industry conferences, and professional collaboration, driving benchmarking initiatives for future hardware procurement.

Qualifications and Experience

  • 5-10 years of experience in HPC systems administration.
  • Demonstrated track record of managing complex HPC systems, including parallel file systems, job schedulers, InfiniBand/Ethernet networks, and monitoring systems.
  • Expertise in supporting users of computational science and engineering, data analysis, and artificial intelligence applications and libraries in different HPC environments.
  • Experience with configuration management tools (Ansible, Puppet, or equivalent).
  • Familiarity with computational science, data analysis, and AI/ML applications and libraries used in HPC environments.
  • Knowledge of project management principles and practices.

Required Competencies

  • Demonstrated ability to support research activities in a highly collaborative HPC environment.
  • Strong analytical, problem-solving, and decision-making skills.
  • Ability to proactively identify and implement system improvements, take initiative, and see tasks through to closure.
  • Ability to manage multiple concurrent projects and deliver high-quality results within deadlines.
  • Proven ability to collaborate cross-functionally with researchers, application teams, and vendors.
  • Effective in multi-cultural, international work environments.
  • Excellent verbal and written communication skills in English, including the ability to prepare and deliver technical reports and presentations.

متطلبات الوظيفة

  • تتطلب ٥-١٠ سنوات خبرة

وظائف مشابهة