Job Title
Senior System Software Engineer – Data Center Compute Diagnostics
Role Summary
Senior engineer responsible for architecting, implementing, debugging, and maintaining low-level diagnostic and stress software that validates next-generation data-center GPUs and rack-scale AI systems. The team develops tests and diagnostics that exercise hardware blocks, interfaces, firmware, drivers, and system-level behavior across validation, manufacturing, qualification, and field support.
Experience Level
Senior — 12+ years of relevant experience expected.
Responsibilities
Hands-on development and technical leadership across diagnostic software and bring-up activities.
- Architect, implement, and maintain diagnostic and stress software in C/C++ and Python for complex hardware systems.
- Lead development efforts, break ambiguous problems into deliverable work, and mentor other engineers.
- Interface with hardware blocks, firmware, Linux device drivers, registers, telemetry, and low-level debugging tools.
- Define diagnostic and stress strategies for engineering validation, manufacturing, product qualification, and field use.
- Design targeted tests for compute engines, memory/cache subsystems, DMA engines, NICs, PCIe/NVLink, power, and thermal behavior.
- Develop workloads from low-level hardware tests to higher-level compute workloads for validation.
- Investigate and resolve complex hardware/software failures across stack layers.
- Collaborate with hardware architects, driver developers, silicon-validation, manufacturing, and field teams to bring up new hardware.
Requirements
Must-have technical skills and experience. Nice-to-have items are listed at the end.
- 12+ years in embedded software, firmware, Linux device drivers, systems software, hardware validation, diagnostics, or silicon bring-up.
- Proven technical leadership of complex software components or projects and experience mentoring engineers.
- Strong programming skills in C and C++; working proficiency in Python.
- Extensive experience developing software that interacts with hardware, firmware, device drivers, and hardware registers.
- Experience with PCIe, NVLink, or networking technologies (Ethernet, InfiniBand) and high-speed interfaces.
- Solid understanding of computer architecture: memory systems, caches, interrupts, DMA, buses, device I/O, and hardware error modes.
- Proven ability to debug complex failures spanning hardware, firmware, drivers, OS, and applications.
- Strong written and verbal communication skills to work across engineering and manufacturing teams.
-
Nice-to-have: prior GPU, CUDA, GEMM-style compute, NCCL, or PyTorch experience.
Education Requirements
BS or MS in Electrical Engineering, Computer Engineering, Computer Science, or a related field, or equivalent practical experience.
About the Company
Company: NVIDIA
Headquarters: Santa Clara, California, USA
NVIDIA is a global leader in accelerated computing, renowned for its innovative solutions in AI and digital twins that transform diverse industries. The company specializes in networking technologies, providing end-to-end InfiniBand and Ethernet solutions for servers and storage that optimize performance and scalability. NVIDIA serves sectors such as high-performance computing, enterprise data centers, and cloud computing, constantly reinventing its products and services to stay ahead in the market.

Date Posted: 2026-07-31