HPC Engineer

Mercedes AMG High Performance Powertrains, Brixworth, GB

About the role

Purpose of the role is to…

- Own the reliability, availability and performance of HPC platforms supporting simulation, analysis and engineering workloads - Improve compute, storage, networking and scheduling services to enable efficient, scalable workload delivery - Provide technical escalation for HPC incidents, capacity issues, performance bottlenecks and complex user problems

THEREFORE WE NEED YOU TO…

Be skilled at…

- Administering Linux-based HPC clusters, including compute nodes, schedulers and shared platform services - Troubleshooting issues across hardware, OS, network, storage, applications and user workflows - Managing capacity, performance and availability for engineering and simulation workloads - Automating operational tasks using Bash, Python, PowerShell, Ansible or equivalent tools - Translating technical user requirements into practical service improvements

Have experience of…

- Supporting Linux-based HPC, scientific computing, simulation or high-throughput compute environments - Diagnosing workload, queue, licence, performance, data movement and application issues - Operating at a senior technical level in an enterprise or engineering-led environment - Delivering maintenance, upgrades, patching and change activity with minimal service impact - Working with suppliers and internal teams to resolve platform issues and improve service maturity

Demonstrate knowledge of…

- HPC architecture, parallel workloads, scheduling, queues and resource allocation - Linux administration, scripting, patching and secure configuration - Schedulers such as Slurm, PBS, LSF or equivalent - Scale-out storage, file systems, backup, archive and data lifecycle management - Networking, interconnects, latency, bandwidth and data locality considerations - Monitoring, performance tuning, benchmarking and capacity forecasting - Security, vulnerability management, access control and compliance for shared platforms - Desirable: motorsport, automotive, CFD, simulation or data science experience

Hold these qualifications…

- Relevant degree, apprenticeship, professional qualification or equivalent technical experience - Relevant technical certifications, or equivalent experience, in Linux, HPC, storage, networking, automation or ITIL

Be…

- Analytical, curious and comfortable solving complex technical problems - Proactive in improving resilience, reducing risk and removing operational friction - Structured, communicative and effective across hands-on delivery and change control - Collaborative, customer-focused and willing to share knowledge

Success in this role will be if you… (deliverables)

- Maintain stable, secure and performant HPC services for critical engineering workloads - Improve compute and storage utilisation through effective monitoring, queue management and capacity planning - Resolve incidents quickly and reduce repeat issues through automation, documentation and service improvement - Deliver upgrades, maintenance and project work safely with clear communication and change control - Improve simulation throughput, data availability and user productivity - Define and guide strategic direction on HPC related topics

Additional information…

- The role combines operational support and project delivery, including planned maintenance, capacity improvement, lifecycle management and occasional out-of-hours activity