System Software Engineer - GPU Power and Performance Management
NVIDIAJob Title
System Software Engineer - GPU Power and Performance Management
Role Summary
Design, implement, and debug production system software that manages GPU power and performance (P-states, DVFS, and power-management controllers) for data-center GPUs. Work across hardware architecture, firmware, device drivers, operating systems, and validation teams through the full product lifecycle from design and pre-silicon development through silicon bring-up and production support.
Experience Level
Entry-level - 2+ years of relevant industry experience, or equivalent internship/co-op/project experience in system software.
Responsibilities
Key duties focused on power/performance control, validation, and cross-team integration.
- Design, implement, and debug GPU power- and performance-management software with emphasis on P-state management and DVFS.
- Develop control policies and mechanisms to optimize GPU performance, power, thermals, and energy efficiency for data-center workloads.
- Support features across the product lifecycle: requirements, architecture, implementation, pre-silicon validation, silicon bring-up, productization, and production support.
- Collaborate with GPU architects and hardware designers to define and implement hardware-software interfaces.
- Analyze interactions among workloads, clocks, voltages, power limits, thermals, telemetry, and system-level policies.
- Investigate, triage, and resolve power, performance, stability, and reliability issues spanning firmware, drivers, hardware, and platform software.
- Execute validation strategies and automation for correctness, transition latency, performance-per-watt, and robustness across conditions.
- Produce clear technical documentation and interface definitions.
Requirements
Must-have technical skills and experience.
- 2+ years relevant industry experience, or equivalent internship/co-op/project experience in system software.
- Strong C programming skills with experience developing and debugging low-level kernel or firmware code.
- Solid understanding of operating-system fundamentals, computer architecture, device-driver architecture, embedded or real-time software, concurrency, and interrupt handling.
- Experience debugging and analyzing issues across multiple software and hardware layers.
- Ability to read and interpret hardware specifications and software interface definitions.
- Strong problem-solving skills and effective written and verbal communication.
Nice-to-have:
- Hands-on or project experience with GPU/CPU/SoC power and performance management, P-states, DVFS, clock/voltage/thermal control, or workload-aware system behavior.
- Experience with validation, scripting/automation, telemetry analysis, and hardware-software interaction testing.
- Experience applying AI tools to software engineering tasks such as coding, validation, or automation workflows.
Education Requirements
BS or MS in Computer Science, Computer Engineering, Electrical Engineering, or a related technical field, or equivalent practical experience. Internship, co-op, or substantial project experience may be considered in lieu of formal degree.
About the Company
Company: NVIDIA
Headquarters: Santa Clara, California, USA
NVIDIA is a global leader in accelerated computing, renowned for its innovative solutions in AI and digital twins that transform diverse industries. The company specializes in networking technologies, providing end-to-end InfiniBand and Ethernet solutions for servers and storage that optimize performance and scalability. NVIDIA serves sectors such as high-performance computing, enterprise data centers, and cloud computing, constantly reinventing its products and services to stay ahead in the market.
