Overview
In this contract role, the AI GPU Infrastructure Architect will play a pivotal role in transforming an existing data centre into a cutting-edge AI-ready facility. The contractor will collaborate with a multidisciplinary team of specialists to shape the design and implementation strategies for large-scale NVIDIA GPU deployments, ensuring robust support for AI training, HPC workloads, and high-density compute environments. This position focuses on architecting solutions that align with AI compute requirements while advocating for efficiency and scalability in a rapidly evolving technological landscape.
Responsibilities
- Define and validate architecture for large-scale NVIDIA GPU deployments.
- Develop AI cluster design standards supporting training and inference workloads.
- Collaborate with data centre design teams to align compute requirements with facility capabilities.
- Define rack density, power, and cooling strategies for high-density AI environments.
- Provide technical leadership around NVIDIA DGX, HGX, Blackwell, and future GPU platforms.
- Design GPU networking and fabric requirements including NVLink, NVSwitch, and InfiniBand.
- Support capacity planning and scaling strategies for a 20-30MW AI compute estate.
- Evaluate liquid cooling and advanced thermal management solutions.
Requirements
- Proven experience designing large-scale GPU infrastructure environments.
- Experience with NVIDIA AI platforms including DGX and/or HGX systems.
- Strong understanding of AI training and inference architectures.
- Deep knowledge of GPU networking technologies including NVLink, NVSwitch, and InfiniBand.
- Knowledge of high-performance compute (HPC) environment design.
- Strong understanding of data centre power and cooling constraints.
- Experience with high-density rack deployments and their challenges.
- Ability to effectively communicate with both technical and executive stakeholders.