Overview
The AI / GPU Infrastructure Architect will play a critical role in designing and deploying large-scale NVIDIA GPU infrastructure tailored for AI and HPC workloads, collaborating with various engineering teams to ensure integration across compute, networking, and storage systems. This position is ideal for a senior architect with specialized experience in configuring high-performance GPU environments, focusing on infrastructure rather than application development or operational support.
Responsibilities
- Design and deploy large-scale NVIDIA GPU/HPC infrastructure.
- Architect multi-node GPU environments across compute, networking, and storage.
- Work with NVIDIA HGX/DGX platforms and high-speed GPU fabrics.
- Commission, validate, and conduct performance testing of GPU clusters.
- Produce high-level designs (HLDs), low-level designs (LLDs), and technical documentation.
- Collaborate with networking, storage, platform, and data-centre engineering teams.
Requirements
- Proven experience as an Infrastructure / Platform Architect with a focus on GPU clusters.
- Hands-on expertise with NVIDIA GPU technologies, including H100/H200.
- Familiarity with high-speed fabrics like InfiniBand and NVLink.
- Experience in building GPU environments using Kubernetes or Slurm.
- Strong background in bare-metal Linux infrastructure.
- Knowledge of automation tools such as Ansible, Terraform, or scripting languages like Python or Bash.