Overview
The GPU Solution Architect will work with a leading AI infrastructure company, leveraging their expertise in GPU architecture to design and deploy scalable solutions for AI and ML workloads. This role involves collaboration with customers, engineering teams, and operations to ensure effective implementation and optimization of high-performance computing systems.
Responsibilities
- Understand customer AI/ML workloads and translate requirements into scalable infrastructure solutions.
- Design GPU clusters covering compute, networking, storage, and capacity planning.
- Work with engineering and operations teams to deliver customer deployments.
- Provide architectural guidance throughout implementation and optimization.
- Support proof of concepts, onboarding, and complex technical deployments.
- Troubleshoot infrastructure and performance challenges.
- Create architecture diagrams, technical documentation, and solution proposals.
- Develop Bills of Materials (BOMs) and support RFQ processes.
Requirements
- 7+ years’ experience in Solutions Architecture, Cloud Architecture, Systems Engineering, or a similar role.
- Strong understanding of GPU infrastructure, AI/ML workloads, and high-performance computing.
- Experience with high-performance networking, including InfiniBand and/or RoCE.
- Strong knowledge of L2/L3 networking, storage, and distributed computing.
- Experience deploying data-center infrastructure and connectivity.
- Hands-on experience with Linux, containers, and orchestration.
- Ability to translate complex customer requirements into practical technical solutions.
- Strong communication and presentation skills for both technical and non-technical audiences.