HPC System Engineer
EDGE Group PJSC · Abu Dhabi, AE · United Arab Emirates · On-site
Sign up free: we match you to jobs like this, tailor your application and fill the form. 2 free applications every day.
ABOUT EDGE
At EDGE bold ideas are engineered into technologies that protect, improve and save lives.
Headquartered in the United Arab Emirates, EDGE is a leading advanced technology group working at the forefront of defence and emerging technologies. Spanning more than 35 specialised companies and multiple centres of excellence, we are purpose-built to move fast. Free from heavy legacy processes, we give our people the autonomy, accountability, and agility to bring breakthrough technologies from concept to reality.
Based in Abu Dhabi, a globally connected hub at the crossroads of Europe, Asia, Africa, and the Middle East, EDGE is home to a truly multicultural community where bold ideas thrive, and the future is shaped.
Together, we are shaping the future.
ABOUT THE ROLE
We are seeking a highly skilled HPC Systems Administrator to manage, optimize, and support our high-performance computing environment used for Computational Fluid Dynamics (CFD) and Finite Element Analysis (FEA) workloads. This role is responsible for the full lifecycle of HPC operations — from infrastructure and scheduler management to user support, performance optimization, and long-term capacity planning. The ideal candidate has strong Linux administration experience, deep knowledge of HPC schedulers, and hands-on familiarity with engineering simulation tools.
RESPONSIBILITIES
A. Infrastructure Management • Maintain and administer compute nodes, login nodes, heterogeneous nodes, NAS storage servers, and high-speed interconnects (InfiniBand). • Manage and maintain HPC-related databases, ensuring timely backups and audit compliance. • Monitor hardware health including CPU temperatures, memory errors, and disk failures. • Ensure high availability, reliability, and minimal downtime across the HPC environment. B. Scheduler & Resource Management • Configure and tune job schedulers (Slurm / PBS / LSF / Grid Engine). • Implement and maintain fair-share scheduling, job priority rules, QoS limits, and preemption policies. • Manage node reservations for large-scale CFD/FEA workloads. • Provide best-practice recommendations for CPU/GPU architecture selection for engineering simulations. • Prevent resource misuse, including node hogging, queue congestion, and starvation of small jobs. C. User Access, Security & Compliance • Manage user accounts, permissions, and storage quotas. • Enforce secure SSH access, MFA, and other security controls. • Ensure compliance with IT governance, data-security, and audit requirements. • Apply OS patches, security updates, and vulnerability fixes. D. Software Stack Management • Install, update, and manage licenses for engineering solvers such as ANSYS, Siemens, NASTRAN, and others. • Maintain version control and apply service pack updates as required. • Manage environment modules (Lmod, Environment Modules). • Optimize compilers, MPI libraries, math libraries, and GPU/graphics-intensive drivers for performance. RESTRICTED E. Performance…