Overview
As an HPC Consultant, you will play a crucial role in improving and maintaining a complex infrastructure environment. Collaborating with the existing Infrastructure team, you will leverage your expertise in Linux and HPC to assess current operational practices, implement enhancements, and ensure reliable day-to-day functionality. This hands-on consulting engagement allows you to drive meaningful improvements while supporting operational needs.
Responsibilities
- Assess and optimise the existing HPC environment, including compute, scheduling, storage, and network performance.
- Provide senior-level Linux administration and troubleshooting across the infrastructure.
- Investigate performance, availability, and capacity issues using monitoring and observability tools.
- Deliver infrastructure improvements through defined work packages and sprint-based activities.
- Improve automation, reliability, and operational processes across the platform.
- Document findings, recommendations, and technical changes for the wider team.
Requirements
- Strong commercial experience administering RHEL/Linux in demanding technical environments.
- Proven HPC experience, ideally with Slurm and distributed compute environments.
- Experience with containers and/or virtualization.
- Strong automation capability using Ansible, Python, and Bash.
- Comfortable diagnosing complex infrastructure and performance problems.
- Able to work independently and communicate technical findings clearly.