AI / GPU Infrastructure Architect

Overview

The AI / GPU Infrastructure Architect will play a critical role in designing and deploying large-scale NVIDIA GPU infrastructure tailored for AI and HPC workloads, collaborating with various engineering teams to ensure integration across compute, networking, and storage systems. This position is ideal for a senior architect with specialized experience in configuring high-performance GPU environments, focusing on infrastructure rather than application development or operational support.

Responsibilities

  • Design and deploy large-scale NVIDIA GPU/HPC infrastructure.
  • Architect multi-node GPU environments across compute, networking, and storage.
  • Work with NVIDIA HGX/DGX platforms and high-speed GPU fabrics.
  • Commission, validate, and conduct performance testing of GPU clusters.
  • Produce high-level designs (HLDs), low-level designs (LLDs), and technical documentation.
  • Collaborate with networking, storage, platform, and data-centre engineering teams.

Requirements

  • Proven experience as an Infrastructure / Platform Architect with a focus on GPU clusters.
  • Hands-on expertise with NVIDIA GPU technologies, including H100/H200.
  • Familiarity with high-speed fabrics like InfiniBand and NVLink.
  • Experience in building GPU environments using Kubernetes or Slurm.
  • Strong background in bare-metal Linux infrastructure.
  • Knowledge of automation tools such as Ansible, Terraform, or scripting languages like Python or Bash.