NVIDIA logo

Principal Software Engineer, GPU Firmware and GPU System Software — CSP Engagements

NVIDIA
August 22, 2026
Full-time
On-site
Santa Clara, California, United States
$272,000 - $431,250 USD yearly
Other Semiconductor Jobs, Level - Senior

Job Title

Principal Software Engineer, GPU Firmware and GPU System Software — CSP Engagements

Role Summary

The Principal Software Engineer will act as NVIDIA's technical focal point for GPU firmware and GPU system software with cloud service provider (CSP) / hyperscale customers. You will work directly with CSP engineering teams to ensure firmware update workflows, recovery procedures, and integration points are reliable at fleet scale and to feed customer-driven priorities into NVIDIA's firmware and system software roadmap.

Experience Level

Senior — 15+ years of experience in GPU system software, firmware, or accelerator platform engineering.

Responsibilities

You will lead technical engagements with CSP customers and drive improvements in firmware manageability, rollout, and operational practices.

  • Lead GPU firmware and system software workstreams with CSP engineering teams; explain firmware architecture, update sequencing, recovery procedures, and power management.
  • Collect and synthesize CSP feedback on manageability, observability, security, and performance; represent those priorities in NVIDIA's roadmap and delivery plans.
  • Design and coordinate firmware update orchestration for large-scale deployments: multi-GPU sequencing, rollback strategies, failure handling, and validation across rack-scale deployments.
  • Serve as the technical interface between NVIDIA and CSP firmware/software engineering to ensure behaviors (error recovery, thermal protection, power transitions) are documented for customer integration.
  • Identify cross-CSP issue patterns and drive documentation, tooling, and test-strategy improvements to reduce fleet-wide operational failures.

Requirements

Must-have technical skills and platform experience; include demonstrated ability to influence customer engineering teams.

  • Deep knowledge of GPU architecture internals and how firmware/driver decisions affect compute performance (SMs, memory hierarchy, GEMM execution, kernel behavior).
  • Experience with multi-GPU fabric architectures and coordination across GPUs in rack-scale systems (e.g., NVLink or similar).
  • Familiarity with GPU firmware components and their interaction with drivers: VBIOS, microcontroller firmware, InfoROM.
  • Experience managing firmware update lifecycles at scale: multi-device sequencing, A/B updates, staged rollouts, rollback, and emergency recovery.
  • Understanding of firmware-level error handling and recovery and how errors propagate to software and applications.
  • Experience with GPU health monitoring and telemetry (Xid errors, thermal/power events, ECC counters) and using telemetry to drive operational decisions.
  • Proven ability to work directly with customers, simplify integration, and influence engineering priorities for improved fleet manageability.

Nice-to-have

  • Direct experience with NVIDIA GPU VBIOS, GPU microcontroller firmware, or GPU driver internals.
  • Experience operating GPU fleets at 10K+ scale (firmware rollout, remediation, fleet configuration management).
  • Experience building runbooks and error taxonomies for GPU firmware behavior and security (secure boot, code signing, attestation, multi-tenancy isolation).
  • Familiarity with GPU power management at fleet scale and its effect on workload performance.

Education Requirements

Bachelor's or Master's degree in Computer Science, Electrical Engineering, or a related technical field, or equivalent practical experience.


About the Company

Company: NVIDIA

Headquarters: Santa Clara, California, USA

NVIDIA is a global leader in accelerated computing, renowned for its innovative solutions in AI and digital twins that transform diverse industries. The company specializes in networking technologies, providing end-to-end InfiniBand and Ethernet solutions for servers and storage that optimize performance and scalability. NVIDIA serves sectors such as high-performance computing, enterprise data centers, and cloud computing, constantly reinventing its products and services to stay ahead in the market.

NVIDIA logo

Date Posted: 2026-08-22