Job Title
Senior Systems Software Engineer - GPU Performance at Scale
Role Summary
Senior engineer responsible for designing and implementing performance tooling, methodologies, and workflows for large-scale GPU-based datacenter and AI compute systems. The role partners with hardware, firmware, software, and customer teams to measure, analyze, and improve performance and stability of AI workloads at scale.
Experience Level
Senior β typically 8+ years of applicable experience; advanced degree (MS/PhD) is desirable.
Responsibilities
Deliver tools, processes, and analyses that validate and improve performance across datacenter products and AI workloads.
- Lead implementation of performance practices and tooling for large-scale GPU infrastructure.
- Align AI workload requirements with datacenter hardware (GPUs, CPUs, networking) and coordinate with internal and customer teams early in the development cycle.
- Build engineering solutions to provide continuous insights into performance regressions and improvements.
- Decompose complex performance or stability issues into minimal, reproducible cases and drive root-cause analysis.
- Collaborate with firmware and software teams (BMC/SBIOS/OS/drivers) to analyze, debug, and resolve critical issues affecting large-scale AI performance.
Requirements
Must-have technical skills and experience.
- Proven understanding of accelerated computing software stacks (CUDA).
- Experience with cloud and container-based enterprise computing architectures; familiarity with Slurm preferred.
- Strong programming and scripting skills in C/C++ and Python; shell scripting (Bash) experience.
- Deep expertise in systems architecture and how components affect performance.
- Experience with container technology and Linux-based OSes (Docker preferred).
- Experience supporting high-performance computing or deep learning workloads in engineering or research environments.
- Strong communication and teamwork skills; results-focused analytical ability.
Nice-to-have:
- End-to-end GPU performance engineering from profiler to system-level analysis.
- Linux systems programming and optimization experience.
- Exposure to virtualization techniques and cloud platform solutions.
- Experience with scheduling and resource management systems and large-scale HPC environments.
Education Requirements
Bachelor's degree in Engineering, Mathematics, Physics, or Computer Science, or equivalent practical experience. Master's or PhD desirable. The posting specifies ~8+ years of applicable experience for senior candidates.
About the Company
Company: NVIDIA
Headquarters: Santa Clara, California, USA
NVIDIA is a global leader in accelerated computing, renowned for its innovative solutions in AI and digital twins that transform diverse industries. The company specializes in networking technologies, providing end-to-end InfiniBand and Ethernet solutions for servers and storage that optimize performance and scalability. NVIDIA serves sectors such as high-performance computing, enterprise data centers, and cloud computing, constantly reinventing its products and services to stay ahead in the market.

Date Posted: 2026-08-21