Platform Architect

Overview

We are seeking an experienced Platform Architect to support a major AI infrastructure program focused on building and scaling NVIDIA AI Factory infrastructure. The architect will play a pivotal role in designing and deploying NVIDIA GPU infrastructure while collaborating closely with teams across infrastructure, networking, storage, security, applications, and data centers. The position is primarily remote, with occasional travel to London required.

Responsibilities

  • Translate complex technical requirements into High-Level Designs (HLDs) and Low-Level Designs (LLDs).
  • Design and deploy NVIDIA GPU infrastructure, specifically focused on HGX GB300, NVL72, and RTX 6000 series servers.
  • Collaborate with infrastructure, networking, storage, security, applications, and data center teams on projects.
  • Utilize Kubernetes for deployment and management of containerized applications.
  • Manage GPU workload scheduling, partitioning, and sharing, incorporating technologies such as MIG and vGPU.
  • Implement infrastructure automation using tools like Terraform and Ansible.
  • Ensure high-performance standards are met in AI fabric environments.

Requirements

  • Proven experience in Platform / Infrastructure Architecture with a focus on compute, storage, networking, and Linux.
  • Expert-level knowledge of Kubernetes architecture.
  • Strong experience with Slurm and Run in GPU/HPC environments.
  • Hands-on experience with NVIDIA HGX GB300 / NVL72 infrastructure, including NVLink, NVSwitch, and Grace Blackwell architecture.
  • Understanding of NVIDIA RTX 6000 series GPU servers and their operational management.
  • Familiarity with InfiniBand, RoCE, Spectrum-X, GPUDirect RDMA/Storage, and related high-performance fabrics.
  • Proficiency in automation and scripting tools such as Terraform, Ansible, Python/shell, Git, and CI/CD methodologies.