HPC Senior Systems Administrator📣 إعلان
| نوع العقد | دوام كامل | |
| طبيعة الوظيفة | بالموقع | |
| الموقع | مكة المكرمة |
وصف الوظيفة
About the Role
KAUST (King Abdullah University of Science and Technology) is seeking a HPC Senior Systems Administrator to join its Supercomputing Laboratory (KSL) in Makkah, Saudi Arabia. This full-time role involves managing a large-scale HPC cluster, its storage systems, and networks, while providing critical support to researchers. The successful candidate will possess 5-10 years of relevant experience in HPC system administration.
Key Responsibilities
- Provide timely user support via multiple channels, maintaining high customer service standards.
- Install, configure, and manage HPC subsystems, including compute nodes, high-performance storage, InfiniBand, Ethernet, and configuration management tools (*, Ansible, Puppet).
- Deploy and manage cluster management software, monitoring tools, and supporting services for HPC clusters.
- Administer the Slurm workload manager, including QOS policies, accounts, accounting, and related automation scripts (Python, C++).
- Develop and maintain automation scripts in Bash and Python to streamline system administration tasks.
- Deploy and manage container environments (Singularity/Apptainer, Docker) for HPC workloads.
- Benchmark HPC system components (CPU, memory, InfiniBand, storage) periodically to ensure optimal performance and identify tuning opportunities.
- Enforce security best practices, including node hardening, kernel patching, and compliance across all systems.
- Manage parallel file systems such as Lustre, GPFS, Weka, or Vast, including performance tuning and capacity planning.
- Directly support research activities in computational science, engineering, data analysis, and AI/ML in collaboration with application support teams.
- Develop software tools and utilities as needed to support research projects on cluster systems.
- Drive proof-of-concept projects and technology evaluations, research industry best practices, and advocate for system enhancements.
- Coordinate with vendors and third-party service providers to report and resolve issues.
- Develop and maintain user documentation, standard operating procedures, and training materials.
Required Qualifications and Experience
- Demonstrated experience (5-10 years) troubleshooting complex hardware issues and documenting root cause analysis.
- Experience managing parallel storage systems (Lustre, GPFS, Weka, Vast, or similar).
- Proven experience benchmarking HPC system components (CPU, memory, InfiniBand, storage).
- Strong Linux system administration experience (RHEL, Rocky Linux, or CentOS) in large-scale HPC environments.
- Experience administering workload managers/schedulers (Slurm, LSF, or PBS).
- Proficiency with configuration management tools such as Ansible or Puppet.
- Familiarity with Kubernetes and container orchestration platforms is desirable.
Essential Skills and Competencies
- Expertise in supporting users of computational science, engineering, data analysis, and artificial intelligence applications and libraries in HPC environments.
- Proficiency with HPC applications and programming models (Fortran, C/C++, Python, MPI, OpenMP, CUDA, OpenACC).
- Demonstrated track record of managing complex HPC systems, including parallel file systems, job schedulers, InfiniBand/Ethernet networks, and monitoring systems.
- Knowledge of project management principles and practices.
- Strong analytical, problem-solving, and decision-making skills.
- Ability to proactively identify and implement system improvements, take initiative, and manage multiple concurrent projects to deliver high-quality results within deadlines.
- Excellent verbal and written communication skills in English, including the ability to prepare and deliver technical reports and presentations.
Work Environment and Collaboration
This role involves close collaboration with faculty, researchers, application support teams, and external partners. The administrator will coordinate with vendors and third-party service providers to resolve complex issues and drive them to closure. Effectiveness in multi-cultural, international work environments and proven cross-functional collaboration skills are essential.
Continuous Development
The successful candidate will be expected to stay at the forefront of HPC advancements through continuous learning, industry conferences, and professional collaboration. This includes driving benchmarking initiatives to inform future hardware procurement and advocating for system enhancements based on industry best practices.
متطلبات الوظيفة
- تتطلب ٥-١٠ سنوات خبرة
وظائف مشابهة
قد يعجبك أيضاً
- وظائف ذات صلة بـ HPC Senior Systems Administrator
- وظائف منسق زهور في الرياض
- وظائف مندوب مبيعات في الرياض
- وظائف موظف استقبال في الرياض
- وظائف مشرف نظافة فندق في الرياض
- وظائف صانع محتوى في الرياض
- مجالات وظيفية أخرى في مكة المكرمة
- وظائف مندوب مبيعات في مكة المكرمة
- وظائف موظف استقبال في مكة المكرمة
- وظائف أخصائي تسويق في مكة المكرمة
- وظائف مدخل بيانات في مكة المكرمة
- وظائف محاسب زبائن (كاشير) في مكة المكرمة
- وظائف Pastry Chef في مكة المكرمة
- وظائف Secretary في مكة المكرمة
- وظائف Barista في مكة المكرمة
- وظائف Human Resources Manager في مكة المكرمة
- وظائف مهندس كهربائي في مكة المكرمة
- استكشف الوظائف في أنحاء المملكة
- وظائف أخصائي تطوير إداري في الرياض
- وظائف مهندس إلكترونيات في مكة المكرمة
- وظائف فني صيانة ميكانيكية في القصيم
- وظائف مهندس ميكانيكي في القصيم
- وظائف مشغل آلات تصنيع منتجات بلاستيكية في الجبيل
- وظائف سائق شاحنة صغيرة في الجفر
- وظائف فني تدفئة وتهوية وتكييف في ضبا
- وظائف مهندس تكاليف في الرياض
- وظائف Reservations Agent في الخبر
- وظائف عامل فرز منتجات في الجبيل