NVIDIA logo

Systems Quality and Reliability Engineer - LPU

NVIDIA
August 17, 2026
Full-time
Remote friendly (Santa Clara, California, United States)
United States
$136,000 - $264,500 USD yearly
Test Engineering Jobs, Level - Mid-Career

Job Title

Systems Quality and Reliability Engineer - LPU

Role Summary

Own and operate failure analysis (FA) and return-material-authorizations (RMA) debug and root-cause processes for NVIDIA AI/ML systems. Lead investigations, produce actionable FA reports, and coordinate resolution with hardware, software, systems, and operations teams.

Work with contract manufacturers (CMs) and internal teams to scale FA capabilities, monitor field quality metrics, and implement containment and mitigation plans.

Experience Level

Mid-level — typically 5+ years of hands-on systems test, validation, or quality/reliability engineering experience.

Responsibilities

Core responsibilities include technical ownership of field failures, FA operations, and quality metrics.

  • Lead debug and root-cause analysis of field RMAs and failed systems; coordinate cross-functional investigations with systems, hardware, software, and operations engineers.
  • Build, manage, and scale RMA and FA processes and capabilities across the organization.
  • Create structured FA reports consistent with 8D or equivalent problem‑solving processes.
  • Analyze RMA/FA/repair data to identify trends; raise quality alerts and drive containment, mitigation, and permanent corrective actions.
  • Monitor hardware quality performance and metrics including RMA rates, MTBF, and reliability ratios.
  • Manage FA operational performance at CMs, ensuring targets for FA cycle time, fault duplication, and fault isolation are met.
  • Oversee setup and qualification of new products into Failure Analysis operations.

Requirements

Must-have technical skills and experience; nice-to-have items listed separately.

  • 5+ years of hands-on systems test, validation, or reliability/quality engineering experience.
  • Proven, practical experience in systems quality and reliability engineering and failure investigation.
  • Hands-on competence with lab instrumentation (oscilloscopes, logic analyzers, power analyzers, etc.).
  • Experience enabling reliability tests (e.g., HTOL) and quality tests (e.g., burn-in).
  • Strong knowledge of fault isolation techniques (examples: OBIRCH, DLS/LADA, LVP, LVI).
  • Proficiency with high-speed interfaces (SerDes, PCIe, DDR) and system/PCB-level test and debug.
  • Programming/scripting proficiency (Python, PERL, C++ or similar) on UNIX/Linux for test automation and data analysis.
  • Ability to manage factory-floor partners and vendor relationships for RMA/FA activities.
  • Nice-to-have: working knowledge of FA tools and techniques such as FIB, SEM, TDR, VNA, and CSAM.

Education Requirements

Bachelor's or Master's degree in Electrical Engineering, Physics, or a related technical field, or equivalent practical experience.


About the Company

Company: NVIDIA

Headquarters: Santa Clara, California, USA

NVIDIA is a global leader in accelerated computing, renowned for its innovative solutions in AI and digital twins that transform diverse industries. The company specializes in networking technologies, providing end-to-end InfiniBand and Ethernet solutions for servers and storage that optimize performance and scalability. NVIDIA serves sectors such as high-performance computing, enterprise data centers, and cloud computing, constantly reinventing its products and services to stay ahead in the market.

NVIDIA logo

Date Posted: 2026-08-17