Senior System Software Architect, HPC and AI Networking
NVIDIAJob Title
Senior System Software Architect, HPC and AI Networking
Role Summary
Architect and prototype scalable system software for distributed AI training and real-time inference, focusing on communication optimization across large-scale HPC and AI deployments. This role partners with framework teams, hardware designers, and customers to improve throughput, latency, and memory efficiency.
Experience Level
Senior — requires substantial hands-on experience; the posting specifies a minimum of 5+ years relevant experience.
Responsibilities
Primary responsibilities include designing, implementing, and evaluating software and protocol improvements for large-scale AI training and inference systems.
- Design and prototype scalable software systems that optimize distributed AI training and inference for throughput, latency, and memory efficiency.
- Develop and evaluate enhancements to communication libraries such as NCCL, UCX, and UCC for deep learning workloads.
- Collaborate with AI framework teams (TensorFlow, PyTorch, JAX) to improve integration, performance, and reliability of communication backends.
- Co-design hardware features in GPUs, DPUs, or interconnects to accelerate data movement and enable new inference/model-serving capabilities.
- Contribute to runtime systems, communication libraries, and AI-specific protocol layers.
- Work with customers to understand deployment needs and deliver practical solutions.
Requirements
Minimum qualifications and core skills required for success in this role.
- 5+ years of experience with deep neural networks, scaling of DNNs, framework parallelism, or deep learning training workloads.
- Deep understanding of inference and training workloads and optimizations (e.g., prefill/decode, data parallelism, tensor parallelism, FDSP).
- Experience with AI network parallelism using collective libraries and RDMA/RoCE.
- Strong background in algorithm design, system programming, and computer architecture.
- Proven programming and software development skills in relevant languages and environments.
- Ability to communicate and collaborate effectively across multinational, multi–time-zone teams.
Nice-to-have
- Experience designing communication middleware for HPC systems, including RoCE and DPU platforms.
- Experience with CUDA and NVIDIA GPU programming models and emerging architectures.
- Demonstrated ability to influence cross-functional engineering and product decisions.
Education Requirements
Ph.D., Master’s, or Bachelor’s degree in Computer Science, Computer Engineering, Electrical Engineering, or a closely related technical field.
About the Company
Company: NVIDIA
Headquarters: Santa Clara, California, USA
NVIDIA is a global leader in accelerated computing, renowned for its innovative solutions in AI and digital twins that transform diverse industries. The company specializes in networking technologies, providing end-to-end InfiniBand and Ethernet solutions for servers and storage that optimize performance and scalability. NVIDIA serves sectors such as high-performance computing, enterprise data centers, and cloud computing, constantly reinventing its products and services to stay ahead in the market.
