Senior Scale-Up Network System Architect
NVIDIAJob Title
Senior Scale-Up Network System Architect
Role Summary
Design the end-to-end NVLink-based scale-up interconnect that links GPUs within and between racks for NVIDIA AI supercomputing platforms. Work at the system level across silicon, firmware, software, and topology teams to convert workload and customer requirements into architecture decisions that will be implemented in production systems.
Experience Level
Senior - 8+ years of industry experience in computer networking, system architecture, or high-performance interconnects.
Responsibilities
Define and validate architecture for next-generation scale-up networks and represent those decisions across engineering teams and customers.
- Define end-to-end system architecture from link and switch behavior through rack- and pod-level topology.
- Translate AI training and inference workload requirements (LLM, MoE, emerging models) into network bandwidth, latency, and resiliency specifications.
- Lead trade-off studies across performance, cost, power, and reliability; build analytical and simulation models to justify decisions.
- Collaborate with ASIC, firmware, software, and systems teams to ensure architectures are implementable in silicon and product.
- Develop and use simulation and analytical models to validate architecture choices prior to silicon commitment.
- Serve as a technical expert on scale-up network behavior in multi-functional and customer-facing discussions.
Requirements
Must-have technical skills and experience.
- 8+ years industry experience in computer networking, system architecture, or high-performance interconnects.
- Deep understanding of network topology, congestion, failure modes, and their impact on distributed workload performance.
- Experience developing or using simulation and modeling environments to evaluate architecture trade-offs.
- Proven ability to drive cross-functional alignment across hardware, firmware, and software teams without direct authority.
- Clear technical communication skills to explain complex architecture trade-offs to technical and non-technical stakeholders.
Nice-to-have:
- Direct experience with NVLink, NVSwitch, or comparable scale-up interconnect technologies.
- Hands-on experience with RoCE/RDMA, low-latency transports, congestion control, and lossless fabric design.
- Experience with memory subsystem architecture (HBM, DDR, LPDDR) and cache coherency.
- Familiarity with large-scale AI model architectures and HPC/supercomputing interconnect development.
Education Requirements
Bachelor's, Master's, or PhD in Computer Science, Computer Engineering, Electrical Engineering, or a related field; or equivalent practical experience.
About the Company
Company: NVIDIA
Headquarters: Santa Clara, California, USA
NVIDIA is a global leader in accelerated computing, renowned for its innovative solutions in AI and digital twins that transform diverse industries. The company specializes in networking technologies, providing end-to-end InfiniBand and Ethernet solutions for servers and storage that optimize performance and scalability. NVIDIA serves sectors such as high-performance computing, enterprise data centers, and cloud computing, constantly reinventing its products and services to stay ahead in the market.
