Platform Architect

Overview

The Platform Architect will play a critical role in the design and deployment of large-scale NVIDIA GPU infrastructure for AI and high-performance computing (HPC) workloads. Collaborating with engineering teams across networking, storage, and data centers, the contractor will focus on building high-performance GPU environments, ensuring the integration of compute, network, and storage components. This position allows for remote work while engaging with a rapidly growing organization at the forefront of GPU technology.

Responsibilities

  • Design and deploy NVIDIA GPU/HPC infrastructure for large-scale projects.
  • Architect multi-node GPU environments integrating compute, network, and storage.
  • Implement high-speed GPU fabrics, including InfiniBand and NVLink.
  • Oversee Kubernetes and/or Slurm-based GPU environments.
  • Manage bare-metal Linux infrastructure for GPU clusters.
  • Commission, validate, and conduct performance testing on GPU clusters.
  • Produce high-level designs (HLDs), low-level designs (LLDs), topology diagrams, and technical documentation.
  • Collaborate with cross-functional teams to ensure cohesive infrastructure development.

Requirements

  • Proven experience designing and deploying NVIDIA GPU clusters.
  • Deep understanding of high-performance computing (HPC) and AI infrastructure.
  • Familiarity with NVIDIA HGX/DGX platforms (H100/H200) and related technologies.
  • Experience with automation tools like Ansible, Terraform, Python, or Bash.
  • Hands-on experience in commissioning and testing GPU clusters.
  • Knowledge of high-speed GPU fabric technologies, including InfiniBand and RoCE.
  • Background in AI Factories, GPU SuperPODs, or NeoCloud infrastructure is preferred.
  • Experience producing technical documentation and architectural designs.