Senior Manager, Failure Analysis
NVIDIAJob Title
Senior Manager, Failure Analysis
Role Summary
Lead NVIDIA's Failure Analysis function, conducting hands-on root-cause investigations of complex hardware, electronic components, and coordinated systems. Build and maintain FA methodologies, lab capabilities, and a team that drives product reliability and corrective actions at scale.
Experience Level
Senior-level. The role requires over 10 years of experience in failure analysis, product reliability, or materials science and at least 5 years of leadership experience.
Responsibilities
Accountable for end-to-end failure analysis, technical leadership, and cross-functional collaboration.
- Lead and perform complex multidisciplinary failure and root-cause investigations using diagnostic tools, lab testing, data analytics, and structured problem solving.
- Design and execute testing programs, laboratory experiments, analytical/computational models, and reproduce field failure conditions.
- Develop, validate, and implement corrective actions and preventive measures to improve product reliability at scale.
- Partner with build, manufacturing, firmware, quality, and supplier teams to drive build changes and resolve cross-functional issues.
- Prepare and present technical reports and presentations for engineering, legal, and business audiences.
- Mentor and develop junior engineers and define failure analysis methodologies, processes, and lab capabilities.
- Identify opportunities to expand FA capabilities and establish strategic partnerships across the supply chain and with external experts.
- Domestic and international travel required (~40%) to support investigations, supplier audits, and client engagements.
Requirements
Must-have technical skills, experience, and leadership competencies.
- Over 10 years of experience in failure analysis, product reliability, or materials science with demonstrated record of leading complex investigations.
- At least 5 years of leadership experience managing technical teams or functions.
- Practical experience with failure analysis techniques and sophisticated lab tools such as SEM, FIB, and digital/RF test instruments.
- Proven ability to troubleshoot hardware/software interactions and reproduce elusive field failure conditions in controlled environments.
- Experience leading multidisciplinary teams and collaborating with engineering, manufacturing, quality, legal, and business stakeholders.
- Strong written and verbal communication skills for conveying complex technical findings to technical and non-technical audiences.
Nice-to-have:
- Experience in the semiconductor, high-performance computing, or AI hardware industry, especially GPU or data-center-scale system reliability.
- Publication record or recognized contributions to failure analysis, materials characterization, or reliability engineering.
- Experience developing lab infrastructure, tool resources, and scaling a failure analysis organization.
Education Requirements
Ph.D. or M.S. in Mechanical Engineering, Chemical Engineering, Aerospace Engineering, Materials Science, or a related field; or equivalent practical experience. (The posting explicitly allows equivalent experience in lieu of the listed degrees.)
About the Company
Company: NVIDIA
Headquarters: Santa Clara, California, USA
NVIDIA is a global leader in accelerated computing, renowned for its innovative solutions in AI and digital twins that transform diverse industries. The company specializes in networking technologies, providing end-to-end InfiniBand and Ethernet solutions for servers and storage that optimize performance and scalability. NVIDIA serves sectors such as high-performance computing, enterprise data centers, and cloud computing, constantly reinventing its products and services to stay ahead in the market.
