System Software Engineer, Performance - CUDA Driver
NVIDIAJob Title
System Software Engineer, Performance - CUDA Driver
Role Summary
Design and ship production C/C++ features and performance optimizations in the CUDA driver and runtime. Work across software and hardware boundaries to analyze workloads, validate platform performance, and deliver solutions that improve latency, throughput, and efficiency for AI, HPC, and graphics workloads.
Experience Level
Mid-level. The role expects relevant systems-software development experience (posting specifies at least 2 years).
Responsibilities
Primary responsibilities focus on performance engineering, feature development, and cross-team technical leadership.
- Design, implement, test, and ship performance-centric features and programming-model capabilities in the CUDA driver and runtime using production-quality C/C++.
- Optimize critical execution paths (kernel launch, synchronization, memory management/movement, CPU–GPU coordination, interconnect) for latency, throughput, bandwidth, efficiency, and scalability.
- Own end-to-end performance problems: characterize workloads, form hypotheses, build measurements and models, isolate root causes across software and hardware, implement solutions, and validate application-level impact.
- Characterize new silicon, establish performance expectations, close software/hardware gaps, and drive performance readiness through product release.
- Translate workload and platform evidence into API, runtime, and systems-software improvements and recommendations for future hardware/architecture.
- Partner with application, library, framework, OS, firmware, GPU architecture, silicon, product, and customer teams; communicate findings and participate in reviews.
- Lead complex feature development and cross-layer investigations; define performance requirements, mentor engineers, and influence hardware/software co-design decisions.
Requirements
Must-have skills and attributes for success in this role.
- Strong production C/C++ systems-programming experience with delivery of substantial features, optimizations, or fixes in a complex codebase.
- Solid operating-systems and concurrency foundations (threads, synchronization, processes, virtual memory, user/kernel interactions).
- Strong computer-architecture knowledge (processors, memory hierarchy, caching/coherence, data movement, system interconnects).
- Proven track record improving software performance: measure behavior, identify bottlenecks, implement fixes, and validate improvements quantitatively.
- Sound technical judgment, ability to own ambiguous problems, and clear communication across teams.
- Direct CUDA or GPU experience is valuable but not required when accompanied by deep systems-software, OS, architecture, and performance-engineering foundations.
Nice-to-have:
- Experience developing GPU or accelerator drivers, runtimes, kernel software, firmware, or compilers.
- Pre-silicon analysis, platform bring-up, performance modeling, or hardware/software co-design experience.
- Systems-level performance experience with AI/DL, HPC, graphics, automotive, or robotics workloads.
- Evidence of technical invention (patents, influential design changes) and scripting skills (Python) for experimentation and analysis.
Education Requirements
BS, MS, or PhD in Computer Science, Computer Engineering, Electrical Engineering, or a related field — or equivalent practical experience.
About the Company
Company: NVIDIA
Headquarters: Santa Clara, California, USA
NVIDIA is a global leader in accelerated computing, renowned for its innovative solutions in AI and digital twins that transform diverse industries. The company specializes in networking technologies, providing end-to-end InfiniBand and Ethernet solutions for servers and storage that optimize performance and scalability. NVIDIA serves sectors such as high-performance computing, enterprise data centers, and cloud computing, constantly reinventing its products and services to stay ahead in the market.
