Overview
The Platform Architect will play a critical role in the design and deployment of large-scale NVIDIA GPU infrastructure for AI and high-performance computing (HPC) workloads. Collaborating with engineering teams across networking, storage, and data centers, the contractor will focus on building high-performance GPU environments, ensuring the integration of compute, network, and storage components. This position allows for remote work while engaging with a rapidly growing organization at the forefront of GPU technology.
Responsibilities
- Design and deploy NVIDIA GPU/HPC infrastructure for large-scale projects.
- Architect multi-node GPU environments integrating compute, network, and storage.
- Implement high-speed GPU fabrics, including InfiniBand and NVLink.
- Oversee Kubernetes and/or Slurm-based GPU environments.
- Manage bare-metal Linux infrastructure for GPU clusters.
- Commission, validate, and conduct performance testing on GPU clusters.
- Produce high-level designs (HLDs), low-level designs (LLDs), topology diagrams, and technical documentation.
- Collaborate with cross-functional teams to ensure cohesive infrastructure development.
Requirements
- Proven experience designing and deploying NVIDIA GPU clusters.
- Deep understanding of high-performance computing (HPC) and AI infrastructure.
- Familiarity with NVIDIA HGX/DGX platforms (H100/H200) and related technologies.
- Experience with automation tools like Ansible, Terraform, Python, or Bash.
- Hands-on experience in commissioning and testing GPU clusters.
- Knowledge of high-speed GPU fabric technologies, including InfiniBand and RoCE.
- Background in AI Factories, GPU SuperPODs, or NeoCloud infrastructure is preferred.
- Experience producing technical documentation and architectural designs.