NVIDIA logo

Principal Software Engineer - Rack Scale Systems Infrastructure

NVIDIA
August 16, 2026
Full-time
On-site
Santa Clara, California, United States
$272,000 - $431,250 USD yearly
Other Semiconductor Jobs, Level - Senior

Job Title

Principal Software Engineer - Rack Scale Systems Infrastructure

Role Summary

Design and lead development of control-plane and infrastructure software for rack-scale products and services where software meets hardware. Work spans orchestration, firmware and OS lifecycle, networking fabrics, and manageability to produce dependable, programmable rack-scale infrastructure for NVIDIA and its customers.

Experience Level

Senior β€” proven experience (15+ years) in systems architecture, system software, distributed systems, or infrastructure engineering.

Responsibilities

Define architecture and lead implementation of infrastructure software that manages rack-scale hardware and fleets.

  • Specify and own software architecture across control plane services, firmware, OS lifecycle, kernel drivers, networking fabrics, and user-mode manageability.
  • Design and build cloud-native and Kubernetes-based controllers, operators, reconciliation loops, and integration APIs for rack and fleet scale.
  • Integrate and align firmware, BMC, BIOS, boot flows, OS images, drivers, and networking with software requirements and roadmaps.
  • Partner with hyperscalers, CSPs, enterprise customers, vendors, and internal teams to validate deployments and reduce production risk.
  • Establish reliability, security, validation, and left-shift strategies for hardware/software delivery.
  • Mentor senior engineers and technical leads; raise engineering standards for large-scale networked systems.
  • Make high-quality technical decisions balancing customer needs, schedule, maintainability, open source adoption, and long-term evolution.

Requirements

Must-have technical skills, plus a concise listing of preferred strengths.

  • 15+ years of professional experience in systems architecture, system software, distributed systems, or infrastructure control planes.
  • Deep architectural knowledge of coordination frameworks: state machines, declarative APIs, reconciliation loops, lifecycle orchestration, failure handling, upgrades, and rollbacks.
  • Production coding ability in Go, C++, or Rust; capable of writing and reviewing production-quality infrastructure software.
  • Experience with Kubernetes or similar orchestration systems as a fabric for managing hardware resources or large-scale infrastructure services.
  • Experience with Linux-based infrastructure, OS rollout and image management, kernel/driver interactions, firmware lifecycle, and hardware bring-up workflows.
  • Strong understanding of data center networking (Ethernet, InfiniBand, RDMA) and fabric-level manageability for accelerator-heavy systems.
  • Knowledge of in-band and out-of-band management architectures (BMCs, Redfish, IPMI) and practical security/update/serviceability tradeoffs.
  • Experience designing software for open source release: API stability, modularity, documentation, and community usability.
  • Strong written and verbal communication; experience specifying requirements and guiding cross-team delivery.
  • Responsible use of AI-assisted development tools to accelerate engineering work.

Nice-to-have:

  • Strong Rust experience for systems and infrastructure software.
  • Hands-on experience with fleet-scale provisioning, updates, rollback, observability, health, and remediation.
  • Experience across full data center product lifecycle (pre/post-silicon, manufacturing, deployment, operations) and open source contribution models.

Education Requirements

BS or MS in Computer Engineering, Computer Science, Electrical Engineering, or a related field, or equivalent practical experience.


About the Company

Company: NVIDIA

Headquarters: Santa Clara, California, USA

NVIDIA is a global leader in accelerated computing, renowned for its innovative solutions in AI and digital twins that transform diverse industries. The company specializes in networking technologies, providing end-to-end InfiniBand and Ethernet solutions for servers and storage that optimize performance and scalability. NVIDIA serves sectors such as high-performance computing, enterprise data centers, and cloud computing, constantly reinventing its products and services to stay ahead in the market.

NVIDIA logo

Date Posted: 2026-08-14