Job Title
Principal Software Engineer, Rack-Scale System Software — CSP Engagements
Role Summary
Technical lead focused on rack-scale system software and firmware for cloud service provider (CSP) engagements. Serve as the primary NVIDIA technical contact for CSP engineering teams to ensure deployable, monitorable, and operable rack-scale systems at fleet scale.
Work across NVIDIA system software, firmware, and CSP partner teams to drive architecture alignment, integrate operational feedback, and ensure integration readiness for fabric management, error recovery, telemetry, and firmware orchestration.
Experience Level
Senior — typically 15+ years of experience in system software, platform firmware, or large-scale distributed systems engineering.
Responsibilities
Own technical engagement with CSPs and drive system-level software architecture, operational readiness, and cross-team execution.
- Align rack-scale SW/FW architecture across CSP engagements: fabric management, link health, NVSwitch/GPU error handling, serviceability, and multi-component firmware orchestration.
- Lead technical work streams with CSP engineering teams to ensure deep understanding of system behavior, recovery policies, and telemetry APIs.
- Capture, synthesize, and champion CSP operational feedback into NVIDIA architecture and product decisions.
- Collaborate with cross-functional teams to translate customer operational requirements into system software and firmware implementation.
- Identify cross-CSP patterns in issues and drive documentation, tooling, and test strategy improvements.
- Enable early customer integration (left-shift) so CSP-side SW/FW integration is completed prior to hardware availability.
- Make technical tradeoff decisions and mitigate execution risks through proactive engagement with CSP partners.
Requirements
Strong background in system-level software, firmware, or distributed systems with demonstrated technical leadership across organizational boundaries.
-
Must-have:
- 15+ years in system software, platform firmware, or large-scale distributed systems engineering.
- Deep understanding of rack-scale software challenges: multi-component coordination, error propagation, health monitoring, and serviceability/reliability.
- Experience with fabric management software, cluster management, or orchestration frameworks; familiarity with firmware architectures and update lifecycle management (sequencing, rollback, recovery).
- Knowledge of error handling and recovery patterns in distributed systems (fault isolation, retry policies, graceful degradation).
- Experience with health monitoring and telemetry systems: health scoring, event correlation, and API design for fleet-level observability.
- Proven ability to provide technical leadership and influence system software design without direct authority; strong communication skills for mentoring customer engineering teams.
-
Nice-to-have:
- Experience with NVIDIA NVSwitch, NVOS, or GPU fabric management software.
- Background in system software for hyperscalers (cluster management, fleet orchestration, health platforms).
- Experience designing error handling and recovery frameworks for hundreds-to-thousands of coordinating devices.
- Familiarity with GPU/accelerator system software (drivers, device/power management) and fleet operations.
- Customer-focused mindset and passion for simplifying CSP operational workflows.
Education Requirements
BS or MS in Computer Science, Electrical Engineering, or a related technical field — or equivalent practical experience.
About the Company
Company: NVIDIA
Headquarters: Santa Clara, California, USA
NVIDIA is a global leader in accelerated computing, renowned for its innovative solutions in AI and digital twins that transform diverse industries. The company specializes in networking technologies, providing end-to-end InfiniBand and Ethernet solutions for servers and storage that optimize performance and scalability. NVIDIA serves sectors such as high-performance computing, enterprise data centers, and cloud computing, constantly reinventing its products and services to stay ahead in the market.

Date Posted: 2026-08-17