Job Title
System Software Engineer β Data Center Compute Diagnostics
Role Summary
Develop low-level diagnostic and stress software for next-generation data-center GPUs and rack-scale AI systems. Work closely with hardware architects, firmware, device drivers, manufacturing, and field teams to bring up and validate new silicon and diagnose system failures.
Focus areas include compute engines, memory/cache subsystems, DMA, PCIe/NVLink, power delivery, and thermal behavior. The role covers design, implementation, validation, productization, and field support of diagnostic components.
Experience Level
Mid-level: typically requires 5+ years of relevant experience in embedded software, firmware, device drivers, systems software, hardware validation, diagnostics, or silicon bring-up.
Responsibilities
Primary responsibilities include developing tests and tools that exercise hardware, analyze failures, and support silicon bring-up.
- Design and implement diagnostic and stress software in C/C++ and Python.
- Integrate tests with firmware, Linux device drivers, hardware registers, telemetry, and low-level debug tools.
- Bring up and validate pre-production silicon and system features.
- Create targeted tests for compute units, memory/cache subsystems, DMA engines, PCIe/NVLink, power delivery, and thermal behavior.
- Investigate hardware and software failures across boundaries (memory errors, ECC, data integrity, performance, thermals, voltage/frequency).
- Contribute to workloads from low-level hardware tests to higher-level AI workloads; develop expertise in CUDA, NCCL, and PyTorch where applicable.
- Provide field support and collaborate with cross-functional teams during productization and deployment.
Requirements
Must-have skills and experience:
- 5+ years of experience in embedded software, firmware, Linux device drivers, systems software, hardware validation, diagnostics, or silicon bring-up.
- Strong programming skills in C and C++; working proficiency in Python.
- Experience developing software that interacts with hardware, firmware, device drivers, registers, or other low-level interfaces.
- Experience creating diagnostics, validation tests, stress tests, or manufacturing tests to isolate hardware or system failures.
- Solid understanding of computer architecture concepts (memory, caches, interrupts, DMA, buses, device I/O).
- Strong debugging and cross-layer problem-solving skills.
- Ability to take ownership of a scoped problem and drive it to completion while collaborating with technical leads and multi-functional teams; good written and verbal communication skills.
- Nice-to-have: prior GPU, CUDA, GEMM, NCCL, or AI workload experience.
Education Requirements
BS or MS in Electrical Engineering, Computer Engineering, Computer Science, or a related field, or equivalent practical experience.
About the Company
Company: NVIDIA
Headquarters: Santa Clara, California, USA
NVIDIA is a global leader in accelerated computing, renowned for its innovative solutions in AI and digital twins that transform diverse industries. The company specializes in networking technologies, providing end-to-end InfiniBand and Ethernet solutions for servers and storage that optimize performance and scalability. NVIDIA serves sectors such as high-performance computing, enterprise data centers, and cloud computing, constantly reinventing its products and services to stay ahead in the market.

Date Posted: 2026-08-01