Senior Deep Learning Frameworks CUDA Software Engineer
NVIDIAJob Title
Senior Deep Learning Frameworks CUDA Software Engineer
Role Summary
Develop and integrate advanced CUDA features and distributed runtime technologies into AI stacks (PyTorch, JAX, TensorRT-related inference engines and similar). Work with teams across hardware, compilers, drivers, and framework maintainers to enable high-performance multi-GPU and multi-node training and inference.
This role focuses on production-quality implementations: prototyping, performance analysis, runtime and compiler interfaces, and scaling solutions for diverse AI workloads.
Experience Level
Senior — typically 8+ years of relevant industry experience (or equivalent academic experience).
Responsibilities
Deliver CUDA and runtime solutions that improve performance, scalability, and reliability of deep learning frameworks and toolchains.
- Integrate new CUDA features and runtime abstractions into AI frameworks from PoC through production.
- Analyze AI workloads and frameworks to identify low-level requirements and optimization opportunities.
- Drive improvements in the compiler-runtime interface for multi-GPU and multi-node execution.
- Design fault-tolerant and elastic runtime solutions for large-scale or dynamic workloads.
- Collaborate with AI researchers, hardware and software architects, compiler and kernel authors, and driver teams to co-design systems.
- Develop profiling and exploratory tooling to measure and accelerate new deep learning paradigms.
- Write clean, maintainable code so prototypes can transition to open-source or production releases.
Requirements
Must-have technical skills and experience for immediate contribution.
- Extensive development experience with deep learning frameworks (PyTorch, JAX) and inference engines (e.g., TRT-LLM, vLLM, sgLang).
- Strong programming and prototyping skills in Python and C++; practical experience with CUDA or related DSLs.
- Proven experience conducting performance benchmarking and using profiler toolchains (PyTorch profiler, NVIDIA Nsight Systems, or similar).
- Solid understanding of AI model parallelisms, distributed ML techniques, and HPC/AI communication concepts.
- Good knowledge of computer system architecture, HW–SW interactions, and operating system fundamentals.
- Ability to work across time zones and collaborate with multiple teams.
Nice-to-have
- Deep expertise in framework internals, autograd, execution graphs, and runtime execution models.
- Hands-on experience with communication libraries (NCCL, MPI, UCX) and distributed training/inference methods (pipeline/tensor parallelism, MoE).
- Experience with deep learning compilers and codegen (Triton, XLA, torch.compile) and kernel authoring.
- Experience optimizing compute and communication overlap in distributed runtimes.
Education Requirements
BS, MS, or PhD in Computer Science, Computer Engineering, Electrical Engineering, or a related technical field; or equivalent practical experience.
About the Company
Company: NVIDIA
Headquarters: Santa Clara, California, USA
NVIDIA is a global leader in accelerated computing, renowned for its innovative solutions in AI and digital twins that transform diverse industries. The company specializes in networking technologies, providing end-to-end InfiniBand and Ethernet solutions for servers and storage that optimize performance and scalability. NVIDIA serves sectors such as high-performance computing, enterprise data centers, and cloud computing, constantly reinventing its products and services to stay ahead in the market.
