Job Title
Senior Software Engineer - NVLink Rack Scale Stability and Reliability
Role Summary
Join the Fabric Networking team focused on NVLink rack-scale systems stability and reliability. The role drives platform bring-up, system validation, diagnostics, and reliability engineering to transition NVLink/NVSwitch platforms into production-ready datacenter deployments.
Experience Level
Senior β typically requires 5+ years of relevant experience in system software, firmware, networking, platform enablement, or data center infrastructure.
Responsibilities
Primary responsibilities center on validation, debugging, automation, and reliability of NVLink-based GPU rack-scale systems.
- Lead platform bring-up, feature enablement, end-to-end software validation, and system debug for NVLink/NVSwitch rack-scale systems.
- Develop tools, diagnostics, automation, and validation infrastructure for regression testing and fleet support.
- Design and execute reliability and MTBI validation via stress tests, telemetry analysis, and failure injection.
- Triage complex issues spanning software, firmware, hardware, networking, and deployment/production environments.
- Collaborate with architecture, hardware, firmware, software, and customer teams to improve system quality and operational readiness.
- Build and maintain SRE-style validation and provisioning systems, monitoring, runbooks, and operational dashboards.
- Create automation and debug workflows to accelerate root-cause analysis and reduce time-to-resolution.
Requirements
Must-have technical skills and experience for successful performance in this role.
Must-have
- 5+ years of experience in system software, firmware, networking, platform enablement, data center infrastructure, or distributed systems.
- Strong programming skills in C/C++ and Python; Bash/Shell scripting is a plus.
- Proven system-level debugging across software, firmware, hardware, and networking layers.
- Solid networking fundamentals: TCP/IP, Ethernet and/or InfiniBand, RDMA/RoCE, routing, switching, and fabric performance analysis.
- Experience with platform bring-up, validation, reliability engineering, stress testing, telemetry analysis, and root-cause debugging for large-scale AI systems.
- Ability to triage multi-domain issues using logs, telemetry, experiments, and structured debugging methods.
- Strong communication and collaboration skills with engineering, customer, and operations teams.
Nice-to-have
- Experience with NVIDIA GPU systems, NVLink, NVSwitch, CUDA, and large-scale AI/HPC clusters.
- Understanding of PCIe, memory hierarchy, DMA, high-speed interconnects, and distributed training/inference systems.
- Experience with server management, cluster provisioning, scaling, fleet monitoring, CI/CD, diagnostics, and reliability tooling.
Education Requirements
BS or MS in Computer Science, Computer Engineering, Electrical Engineering, or a related technical field; or equivalent practical experience.
About the Company
Company: NVIDIA
Headquarters: Santa Clara, California, USA
NVIDIA is a global leader in accelerated computing, renowned for its innovative solutions in AI and digital twins that transform diverse industries. The company specializes in networking technologies, providing end-to-end InfiniBand and Ethernet solutions for servers and storage that optimize performance and scalability. NVIDIA serves sectors such as high-performance computing, enterprise data centers, and cloud computing, constantly reinventing its products and services to stay ahead in the market.

Date Posted: 2026-08-14