HPC Consultant

Overview

As an HPC Consultant, you will play a crucial role in improving and maintaining a complex infrastructure environment. Collaborating with the existing Infrastructure team, you will leverage your expertise in Linux and HPC to assess current operational practices, implement enhancements, and ensure reliable day-to-day functionality. This hands-on consulting engagement allows you to drive meaningful improvements while supporting operational needs.

Responsibilities

  • Assess and optimise the existing HPC environment, including compute, scheduling, storage, and network performance.
  • Provide senior-level Linux administration and troubleshooting across the infrastructure.
  • Investigate performance, availability, and capacity issues using monitoring and observability tools.
  • Deliver infrastructure improvements through defined work packages and sprint-based activities.
  • Improve automation, reliability, and operational processes across the platform.
  • Document findings, recommendations, and technical changes for the wider team.

Requirements

  • Strong commercial experience administering RHEL/Linux in demanding technical environments.
  • Proven HPC experience, ideally with Slurm and distributed compute environments.
  • Experience with containers and/or virtualization.
  • Strong automation capability using Ansible, Python, and Bash.
  • Comfortable diagnosing complex infrastructure and performance problems.
  • Able to work independently and communicate technical findings clearly.