NVIDIA logo

Systems Quality and Reliability Engineer - LPU

NVIDIA
August 16, 2026
Full-time
Remote friendly (Santa Clara, California, United States)
Worldwide
$136,000 - $264,500 USD yearly
Test Engineering Jobs, Level - Mid-Career

Job Title

Systems Quality and Reliability Engineer - LPU

Role Summary

Own, build, and manage RMA and failure-analysis (FA) debug and root-cause efforts for NVIDIA AI/ML products in the LPU team. Work cross-functionally with systems, hardware, software, and operations teams to identify and mitigate field quality issues.

NVIDIA develops advanced GPU and AI systems; this role supports product reliability and field returns at scale.

Experience Level

Mid-level. The role expects 5+ years of hands-on systems test, validation, or reliability engineering experience.

Responsibilities

Primary responsibilities include ownership of RMA/FA processes, quality metrics, and coordination with manufacturing partners.

  • Lead debug and root-cause analysis of field RMAs and FAs; coordinate cross-functional investigations.
  • Scale FA capabilities and standardize FA processes (8D or similar reporting).
  • Analyze RMA, FA and repair data to identify trends; raise quality alerts and drive containment and mitigation plans.
  • Monitor hardware quality metrics including RMA rates, MTBF, and reliability ratios.
  • Manage FA operational performance at contract manufacturers (CMs): cycle times, fault duplication, and fault isolation rates.
  • Oversee setup and ramp of new products into Failure Analysis operations.

Requirements

Must-have skills and experience:

  • 5+ years hands-on systems test, validation, or reliability engineering experience.
  • Proven practical experience in systems quality and reliability engineering.
  • Proficiency with lab equipment such as oscilloscopes, logic analyzers, and power analyzers.
  • Experience enabling reliability tests (HTOL) and quality tests (Burn-in).
  • Strong knowledge of fault isolation techniques (OBIRCH, DLS/LADA, LVP, LVI).
  • Proficiency with high-speed interfaces (SerDes, PCIe, DDR).
  • Proficiency in scripting/programming (Python, Perl, C++) on UNIX/Linux.
  • Experience with PCB card and system-level test and debug; ability to manage factory/CM partners for RMA/FA activities.

Nice-to-have:

  • Working knowledge of FA techniques and tools such as FIB, SEM, TDR, VNA, and CSAM.

Education Requirements

BS or MS in Electrical Engineering, Physics, or a related technical field, or equivalent practical experience.


About the Company

Company: NVIDIA

Headquarters: Santa Clara, California, USA

NVIDIA is a global leader in accelerated computing, renowned for its innovative solutions in AI and digital twins that transform diverse industries. The company specializes in networking technologies, providing end-to-end InfiniBand and Ethernet solutions for servers and storage that optimize performance and scalability. NVIDIA serves sectors such as high-performance computing, enterprise data centers, and cloud computing, constantly reinventing its products and services to stay ahead in the market.

NVIDIA logo

Date Posted: 2026-08-14