Job Title
Principal Firmware Engineer – Server Manageability and Observability
Role Summary
Architect and own system software architecture for NVIDIA data center platforms (DGX, HGX), spanning firmware, kernel drivers, operating systems, and user-mode components. Collaborate with internal component leads and hyperscaler customers to define requirements, KPIs, and delivery of next-generation data center products.
Experience Level
Senior level: requires extensive experience; this role specifies 15+ years in system architecture and design.
Responsibilities
Primary responsibilities include technical leadership, customer engagement, and cross-functional system architecture.
- Serve as primary technical contact for major customers: lead technical discussions, define KPIs, gather requirements, and resolve complex technical issues.
- Lead system software architecture and technical strategy for next-generation data center products, including firmware and kernel-level design.
- Drive strategic collaborations with hyperscalers and align NVIDIA roadmaps to customer requirements.
- Develop and promote adoption of new technologies and protocols across platforms.
- Make critical technical decisions in ambiguous situations and mitigate program risk via left-shift strategies.
- Coordinate cross-functional teams to deliver complex, large-scale projects to completion.
Requirements
Must-have technical skills, experience, and behaviours for success.
- Deep expertise in scalable, high-performance server system architecture with emphasis on SW/HW interfaces.
- Extensive experience with complex system software for accelerators (GPUs, DPUs, FPGAs).
- Mastery of system firmware (SBIOS, OpenBMC), embedded systems, and Linux kernel internals.
- Proficiency in out-of-band and in-band management architectures and device protocols such as MCTP, PLDM, SPDM, RDE, and system management APIs like Redfish and IPMI.
- Strong knowledge of networking technologies and protocols (TCP/IP, Ethernet, InfiniBand) and advanced switching/routing concepts.
- Experience collaborating with platform security teams to balance security and usability trade-offs.
- Proven record leading complex, cross-functional programs and applying left-shift risk mitigation.
- 15+ years of experience in system architecture and design.
Nice-to-have:
- Experience with cloud and cluster deployment/management systems; contributions to standards bodies (OCP, DMTF).
- Familiarity with NVIDIA HPC programming models and libraries (CUDA, cuDNN, DOCA).
- Knowledge of enterprise storage architectures and distributed parallel processing paradigms.
Education Requirements
BS or MS in Computer Science, Electrical Engineering, or a related technical field, or equivalent practical experience.
About the Company
Company: NVIDIA
Headquarters: Santa Clara, California, USA
NVIDIA is a global leader in accelerated computing, renowned for its innovative solutions in AI and digital twins that transform diverse industries. The company specializes in networking technologies, providing end-to-end InfiniBand and Ethernet solutions for servers and storage that optimize performance and scalability. NVIDIA serves sectors such as high-performance computing, enterprise data centers, and cloud computing, constantly reinventing its products and services to stay ahead in the market.

Date Posted: 2026-08-17