Senior System Software Engineer – Data Center Compute Diagnostics
NVIDIAJob Title
Senior System Software Engineer – Data Center Compute Diagnostics
Role Summary
Lead design and implementation of low-level diagnostic and stress software that exercises and validates data center GPUs and rack-scale AI systems. The team develops tools that interact directly with hardware, firmware, device drivers, registers, and telemetry to bring up new silicon and diagnose system failures.
Experience Level
Senior — typically requires 12+ years of experience in embedded software, firmware, device drivers, systems software, hardware validation, or silicon bring-up.
Responsibilities
Hands-on development, technical leadership, and cross-team collaboration to deliver diagnostic software from validation through productization and field support.
- Architect and implement diagnostic and stress software in C/C++ and Python for complex hardware systems.
- Lead multi-engineer development efforts: decompose ambiguous problems, prioritize work, and mentor engineers.
- Interface with hardware blocks, firmware, Linux device drivers, hardware registers, telemetry, and low-level debug tools.
- Define diagnostic and stress strategies for engineering validation, manufacturing, qualification, and field use.
- Design and run targeted tests for compute engines, memory/cache subsystems, DMA engines, NICs, PCIe/NVLink, power, and thermal behavior.
- Develop workloads from low-level GPU tests to higher-level AI workloads (CUDA, GEMM-style compute, NCCL, PyTorch).
- Investigate complex failures across hardware, firmware, drivers, OS, and applications to determine root causes.
- Use modern development and analysis tools (including AI-assisted tools where appropriate) to accelerate coding, debugging, and failure analysis.
Requirements
Must-have technical skills, leadership experience, and hands-on low-level development background.
- 12+ years experience in embedded software, firmware, Linux device drivers, systems software, hardware validation, diagnostics, or silicon bring-up.
- Proven technical leadership on complex software components or projects; experience coordinating cross-functional teams and mentoring engineers.
- Strong programming skills in C and C++; working proficiency in Python.
- Extensive experience developing software that interacts with hardware, firmware, device drivers, hardware registers, or other low-level interfaces.
- Experience debugging complex failures spanning hardware, firmware, drivers, operating systems, and applications.
- Background with PCIe, NVLink, or networking technologies (Ethernet, InfiniBand) and strong understanding of computer architecture (memory systems, caches, interrupts, DMA, buses, device I/O).
- Ability to define technical direction, make engineering tradeoffs, and drive ambiguous problems to completion across organizational boundaries.
- Excellent written and verbal communication skills for interaction with architects, manufacturing, field engineers, and leadership.
- Nice-to-have: prior GPU/CUDA/GEMM experience and familiarity with NCCL or PyTorch.
Education Requirements
BS or MS in Electrical Engineering, Computer Engineering, Computer Science, or a related field, or equivalent practical experience.
About the Company
Company: NVIDIA
Headquarters: Santa Clara, California, USA
NVIDIA is a global leader in accelerated computing, renowned for its innovative solutions in AI and digital twins that transform diverse industries. The company specializes in networking technologies, providing end-to-end InfiniBand and Ethernet solutions for servers and storage that optimize performance and scalability. NVIDIA serves sectors such as high-performance computing, enterprise data centers, and cloud computing, constantly reinventing its products and services to stay ahead in the market.
