Overview
We are seeking a skilled GPU / AI Network Engineer to design and manage high-performance data center network infrastructures. This remote role focuses on creating advanced networking solutions for AI, GPU, and high-performance computing environments. The engineer will collaborate with senior infrastructure and engineering teams to develop networking strategies that meet the demands of large-scale and mission-critical services.
Responsibilities
- Design scalable Leaf-Spine architectures supporting AI and HPC workloads.
- Build and optimize 100G, 400G, and next-generation 800G networks.
- Deliver advanced routing and fabric technologies, including BGP EVPN/VXLAN and high-availability architectures.
- Architect AI network fabrics using RoCE, RDMA, InfiniBand, and lossless Ethernet.
- Lead new data center deployments and carrier onboarding processes.
- Design secure and resilient connectivity solutions for Azure, AWS, and Google Cloud.
- Drive operational excellence across monitoring, capacity planning, and disaster recovery efforts.
Requirements
- Proven experience in network architecture for AI Infrastructure, GPU Compute Environments, and HPC Platforms.
- Strong expertise in Leaf-Spine Networks, BGP EVPN/VXLAN, and Routing & Switching.
- Experience with network security and carrier connectivity solutions.
- Familiarity with high-performance Ethernet technologies and 800G Ethernet deployments.
- Knowledge of RoCE, RDMA, and InfiniBand architecture.
- Understanding of optical networking, dark fiber, and wavelength services.
- Experience with data center commissioning and multi-carrier resilience strategies.