Senior Software Development Engineer in Test - Datacenter Server OS
NVIDIAJob Title
Senior Software Development Engineer in Test - Datacenter Server OS
Role Summary
Join NVIDIA's platform SWQA team to develop and execute reliability and validation tests for datacenter server platforms (HGX/DGX/MGX). The role focuses on server, OS, firmware and CUDA software stack testing, automation, CI/CD, and large-scale cluster validation.
Collaborate with cross-functional engineering and supplier teams to design test plans, build automation frameworks, diagnose failures, and improve platform reliability for AI and HPC deployments.
Experience Level
Senior-level. Preferred 5+ years of relevant experience; a master's degree is considered an alternative level of qualification.
Responsibilities
Primary responsibilities include platform validation, automation development, and driving resolution of reliability issues.
- Develop and execute platform test plans for servers, OS, firmware, and CUDA software stacks.
- Install, configure, and validate operating systems, firmware, and server software on bare-metal and virtualized systems.
- Design, build, and maintain front-end and back-end automation frameworks and automated tests.
- Perform root-cause analysis on reliability and validation failures and implement mitigations.
- Review partner and supplier test results and recommend additional component or system-level testing as needed.
- Manage defect lifecycles and coordinate with cross-functional teams to drive solutions.
- Work within an agile software development team maintaining high production quality standards.
Requirements
Required technical skills and proven experience; education details are listed below.
- Proven experience with server- and OS-level automation, CI/CD pipelines, and DevOps processes.
- Strong scripting/programming skills (Python, Shell); experience with Ansible and Jenkins.
- Familiarity with C/C++, Java, or JavaScript for test tooling and automation.
- Deep Linux troubleshooting and debugging experience across distributions (Ubuntu, Red Hat, CentOS, SuSE, Fedora).
- Experience with bare-metal and virtualized environments (KVM, VMware, Hyper-V) and PXE deployment.
- Experience using AI frameworks and tools for test-plan creation, test-case development, and benchmarking (e.g., TensorFlow, PyTorch).
- Familiarity with Git workflows (GitHub/GitLab/Gerrit) and container/orchestration tools (Docker, Kubernetes, SLURM) is beneficial.
Nice-to-have
- Experience with firmware, BMC/OpenBMC, Redfish, PCIe, enterprise storage, networking, and low-level system interfaces.
- Experience working with NVIDIA GPU hardware, CUDA/OpenCL, LLM/NLP benchmarking, and AI-related tooling.
- Background in parallel programming and large-scale cluster validation.
Education Requirements
Bachelor's degree in a STEM field (Science, Technology, Engineering, Mathematics, or Physics) or equivalent practical experience. A master's degree is referenced as an alternate qualification.
About the Company
Company: NVIDIA
Headquarters: Santa Clara, California, USA
NVIDIA is a global leader in accelerated computing, renowned for its innovative solutions in AI and digital twins that transform diverse industries. The company specializes in networking technologies, providing end-to-end InfiniBand and Ethernet solutions for servers and storage that optimize performance and scalability. NVIDIA serves sectors such as high-performance computing, enterprise data centers, and cloud computing, constantly reinventing its products and services to stay ahead in the market.
