HPC Linux Engineer

Overview

The HPC Linux Engineer will play a critical role in optimizing and enhancing an existing HPC and Linux infrastructure within a dynamic and technically ambitious team. This hands-on position involves direct engagement with day-to-day operations as well as spearheading improvement initiatives focused on the computational platform. The engineer will collaborate with the existing infrastructure team, ensuring reliable operations while implementing practical enhancements based on thorough assessments of current processes.

Responsibilities

  • Assess and optimise the existing HPC environment, including compute, scheduling, storage, and network performance.
  • Provide senior-level Linux administration and troubleshooting across the infrastructure.
  • Investigate performance, availability, and capacity issues using monitoring and observability tooling.
  • Deliver infrastructure improvements through defined work packages and sprint-based activity.
  • Improve automation, reliability, and operational processes across the platform.
  • Document findings, recommendations, and technical changes for the wider team.

Requirements

  • Strong commercial experience administering RHEL/Linux in demanding technical environments.
  • Proven HPC experience, ideally working with Slurm and distributed compute environments.
  • Experience with containers and/or virtualisation.
  • Strong automation capability using Ansible, Python, and Bash.
  • Comfortable diagnosing complex infrastructure and performance problems.
  • Able to work independently and communicate technical findings clearly.