Job Title
Senior Software Development Engineer in Test - Datacenter Server OS
Role Summary
Member of the platform SWQA team responsible for developing and executing reliability and validation test plans for NVIDIA HGX/DGX/MGX servers. Focus areas include server OS, firmware, CUDA software stack, automation frameworks, and cluster-scale reliability testing.
The role drives automation, CI/CD, root-cause analysis of failures, and cross-team coordination to improve hardware and software platform quality at scale.
Experience Level
Senior β typically requires 5+ years of relevant industry experience.
Responsibilities
Primary responsibilities include test planning, execution, automation, and failure analysis for datacenter servers and software stacks.
- Develop and execute platform test plans for servers, OS, firmware, and CUDA software stacks.
- Install, configure, and validate system OS, firmware, and server software on bare-metal and virtual environments.
- Design, build, and maintain front-end and back-end automation frameworks and tests.
- Drive root-cause analysis for reliability and validation test failures and implement mitigations.
- Manage bug lifecycle and collaborate across hardware, firmware, and software teams.
- Review partner and supplier test results and recommend additional reliability testing when needed.
- Work within an agile development team and contribute to CI/CD and DevOps practices.
Requirements
Key technical skills and experience required or strongly preferred. Items labelled "Must-have" are essential; "Nice-to-have" items improve candidacy.
-
Must-have: 5+ years of experience in OS and server-level automation, CI/CD, and DevOps workflows.
-
Must-have: Strong scripting and programming skills (Python, Shell); experience with automation tools such as Ansible and CI systems like Jenkins.
-
Must-have: Deep Linux troubleshooting and debugging on Ubuntu/RedHat/CentOS/SuSE/Fedora in bare-metal and KVM/VMWare/Hyper-V environments.
-
Must-have: Experience building and running automated test frameworks, and using source-control systems (GitHub/GitLab/Gerrit).
-
Must-have: Experience creating test plans and automating test cases; familiarity with telemetry and reliability metrics.
-
Nice-to-have: Familiarity with AI tools/frameworks (TensorFlow, PyTorch) and using AI tooling for test-plan and test-case generation.
-
Nice-to-have: Experience with firmware/BMC/OpenBMC, Redfish, PCIe, storage, network protocols, PXE, SLURM, Kubernetes, Docker, and cluster orchestration.
-
Nice-to-have: Background in CUDA/OpenCL or other parallel programming models and direct experience with NVIDIA GPU hardware.
Education Requirements
Bachelor's degree in a STEM field (Science, Technology, Engineering, Mathematics, or Physics) or equivalent practical experience. A master's degree is considered equivalent to increased experience for this role. Certifications or formal credentials are not required but may be beneficial.
About the Company
Company: NVIDIA
Headquarters: Santa Clara, California, USA
NVIDIA is a global leader in accelerated computing, renowned for its innovative solutions in AI and digital twins that transform diverse industries. The company specializes in networking technologies, providing end-to-end InfiniBand and Ethernet solutions for servers and storage that optimize performance and scalability. NVIDIA serves sectors such as high-performance computing, enterprise data centers, and cloud computing, constantly reinventing its products and services to stay ahead in the market.

Date Posted: 2026-08-21