Lead Systems Software Test Engineer – CSP Engagements
NVIDIAJob Title
Lead Systems Software Test Engineer – CSP Engagements
Role Summary
Lead validation and test engineering for NVIDIA datacenter rack-scale products with a focus on Cloud Service Provider (CSP) integrations. The role combines full-stack systems validation (cluster-to-rack) with customer-facing activities to ensure stable, performant ML training and inference platforms from conception through deployment.
Experience Level
Senior-level. Requires approximately 8+ years of systems software validation or equivalent experience.
Responsibilities
Work across engineering teams and CSP partners to define test strategy, reproduce and triage issues, and confirm release readiness.
- Define test strategies and validation plans for CSP integration milestones; align with partner test methodologies and provide recommendations.
- Reproduce, characterize, and triage customer bugs in partner environments; produce summary test reports for releases and NPI phases.
- Validate fixes, mitigations, and release updates against deployed CSP software and partner configurations.
- Drive root-cause analysis with NVIDIA development teams and provide clear pass/fail evidence for release readiness.
- Collaborate with CSP teams on provisioning, access, break-fix workflows, and environment readiness; produce concise release-readiness summaries.
- Manage large test output datasets and develop tooling for debug data retrieval, visualization, and reporting.
- Localize problems with targeted reproduction steps, stress and edge-case testing, and replicate issues in the lab.
- Run performance benchmarks for training and inference and collaborate with AE/FAE/Solution Architect teams on validations and documentation.
Requirements
Must-have technical skills and experience for successful performance in this role.
- 8+ years of experience in validation, QA, system test, diagnostics, platform bring-up, or release qualification for complex HW–SW systems.
- Strong understanding of server platforms, firmware, drivers, OS integration, networking, and large-scale cluster environments.
- Hands-on debugging across hardware, firmware, software, networking, and infrastructure layers; ability to analyze logs, telemetry, diagnostics, and automation failures.
- Familiarity with Linux, shell scripting, Python (or similar) for automation, and CI/regression workflows; experience creating test plans, regression suites, and defect documentation.
- Proficient in Python with test automation and test-infrastructure design experience.
- Strong cross-functional communication skills with QA, development, field, support, and customer engineering teams.
- Customer-facing experience: partner collaboration, on-site or remote troubleshooting, and publishing clear engineering summaries.
- Nice-to-have: cloud and cluster deployment experience, MLOps familiarity, and experience running deep learning workloads and related automation.
Education Requirements
BS or MS in Computer Engineering, Computer Science, or a related field — or equivalent practical experience.
About the Company
Company: NVIDIA
Headquarters: Santa Clara, California, USA
NVIDIA is a global leader in accelerated computing, renowned for its innovative solutions in AI and digital twins that transform diverse industries. The company specializes in networking technologies, providing end-to-end InfiniBand and Ethernet solutions for servers and storage that optimize performance and scalability. NVIDIA serves sectors such as high-performance computing, enterprise data centers, and cloud computing, constantly reinventing its products and services to stay ahead in the market.
