img
نوع العقددوام كامل
طبيعة الوظيفةبالموقع
الموقعمكة المكرمة

وصف الوظيفة

About the Role

KAUST (King Abdullah University of Science and Technology) is seeking a highly motivated and skilled HPC Senior Systems Administrator to join the KAUST Supercomputing Laboratory (KSL). This full-time position is based in Makkah, Saudi Arabia, and requires 5-10 years of experience. The successful candidate will be responsible for managing an HPC cluster of approximately 600 CPU and GPU nodes, HPC storage systems, InfiniBand and Ethernet networks, and day-to-day operational issues. The role provides broad support to researchers and end-users across computational science, engineering, big data analysis, and artificial intelligence/machine learning workloads.

Key Responsibilities

  • Provide timely and effective user support via telephone, walk-in, email, and ticketing system for all inquiry types, maintaining high customer service standards.
  • Install, configure, and manage HPC subsystems including compute nodes, high-performance storage systems, InfiniBand, Ethernet, and configuration management tools (*, Ansible, Puppet).
  • Deploy and manage cluster management software, monitoring tools, and supporting services for operating HPC clusters.
  • Install and administer the Slurm workload manager, manage QOS policies, accounts, accounting, and related automation scripts (Python and C++).
  • Develop and maintain automation scripts in Bash and Python to streamline system administration tasks.
  • Deploy and manage container environments (Singularity/Apptainer, Docker) for HPC workloads.
  • Benchmark HPC system components like CPU, memory, InfiniBand, and storage periodically to ensure optimal performance and identify tuning opportunities across hardware, driver, and application layers.
  • Enforce security best practices including node hardening, kernel patching, and compliance across all systems.
  • Manage parallel file systems such as Lustre, GPFS, Weka, or Vast, including performance tuning and capacity planning.
  • Directly support research activities in computational science, engineering, data analysis, and AI/ML by working closely with faculty, researchers, collaboration partners, and industrial partners in collaboration with application support teams.
  • Develop software tools and utilities as needed to support research projects on cluster systems and subsystems.
  • Drive proof-of-concept projects and technology evaluations end-to-end, research industry best practices, and advocate system enhancements.
  • Coordinate with vendors and third-party service providers to report and resolve issues in a timely manner.
  • Develop and maintain user documentation, standard operating procedures, and training materials in the internal wiki.
  • Stay at the forefront of HPC advancements through continuous learning, industry conferences, and professional collaboration, while driving benchmarking initiatives to inform future hardware procurement.

Required Competencies

  • Expertise in supporting users of computational science and engineering, data analysis, and artificial intelligence applications and libraries in different HPC environments.
  • Strong expertise in Linux system administration (RHEL, Rocky Linux, or CentOS) in large-scale HPC environments.
  • Proficiency with HPC applications and programming models (Fortran, C/C++, Python, MPI, OpenMP, CUDA, OpenACC).
  • Demonstrated track record of managing complex HPC systems, including parallel file systems, job schedulers, InfiniBand/Ethernet networks, and monitoring systems.
  • Experience with configuration management tools (Ansible, Puppet, or equivalent).
  • Familiarity with computational science, data analysis, and AI/ML applications and libraries used in HPC environments.
  • Knowledge of project management principles and practices.
  • Demonstrated ability to support research activities in a highly collaborative HPC environment.

Skills and Attributes

  • Strong analytical, problem-solving, and decision-making skills.
  • Proactively identifies and implements system improvements; takes initiative and sees tasks through to closure.
  • Ability to manage multiple concurrent projects and deliver high-quality results within deadlines.
  • Proven ability to collaborate cross-functionally with researchers, application teams, and vendors.
  • Effective in multi-cultural, international work environments.
  • Excellent verbal and written communication skills in English, including the ability to prepare and deliver technical reports and presentations.

متطلبات الوظيفة

  • تتطلب ٥-١٠ سنوات خبرة

وظائف مشابهة