HPC Engineer
Mercedes AMG High Performance Powertrains, Brixworth, GB
About the role
Purpose of the role is to…
- Own the reliability, availability and performance of HPC platforms supporting simulation, analysis and engineering workloads - Improve compute, storage, networking and scheduling services to enable efficient, scalable workload delivery - Provide technical escalation for HPC incidents, capacity issues, performance bottlenecks and complex user problems
THEREFORE WE NEED YOU TO…
Be skilled at…
- Administering Linux-based HPC clusters, including compute nodes, schedulers and shared platform services - Troubleshooting issues across hardware, OS, network, storage, applications and user workflows - Managing capacity, performance and availability for engineering and simulation workloads - Automating operational tasks using Bash, Python, PowerShell, Ansible or equivalent tools - Translating technical user requirements into practical service improvements
Have experience of…
- Supporting Linux-based HPC, scientific computing, simulation or high-throughput compute environments - Diagnosing workload, queue, licence, performance, data movement and application issues - Operating at a senior technical level in an enterprise or engineering-led environment - Delivering maintenance, upgrades, patching and change activity with minimal service impact - Working with suppliers and internal teams to resolve platform issues and improve service maturity
Demonstrate knowledge of…
- HPC architecture, parallel workloads, scheduling, queues and resource allocation - Linux administration, scripting, patching and secure configuration - Schedulers such as Slurm, PBS, LSF or equivalent - Scale-out storage, file systems, backup, archive and data lifecycle management - Networking, interconnects, latency, bandwidth and data locality considerations - Monitoring, performance tuning, benchmarking and capacity forecasting - Security, vulnerability management, access control and compliance for shared platforms - Desirable: motorsport, automotive, CFD, simulation or data science experience
Hold these qualifications…
- Relevant degree, apprenticeship, professional qualification or equivalent technical experience - Relevant technical certifications, or equivalent experience, in Linux, HPC, storage, networking, automation or ITIL
Be…
- Analytical, curious and comfortable solving complex technical problems - Proactive in improving resilience, reducing risk and removing operational friction - Structured, communicative and effective across hands-on delivery and change control - Collaborative, customer-focused and willing to share knowledge
Success in this role will be if you… (deliverables)
- Maintain stable, secure and performant HPC services for critical engineering workloads - Improve compute and storage utilisation through effective monitoring, queue management and capacity planning - Resolve incidents quickly and reduce repeat issues through automation, documentation and service improvement - Deliver upgrades, maintenance and project work safely with clear communication and change control - Improve simulation throughput, data availability and user productivity - Define and guide strategic direction on HPC related topics
Additional information…
- The role combines operational support and project delivery, including planned maintenance, capacity improvement, lifecycle management and occasional out-of-hours activity