Server Performance Architect - Hardware
NVIDIAJob Title
Server Performance Architect - Hardware
Role Summary
The Server Performance Architect will define and drive system-level performance targets for next-generation AI server platforms. The role focuses on hands-on workload characterization, system-level profiling, and cross-functional collaboration to close performance gaps from bring-up through production.
Work spans CPU, GPU, memory, interconnect, networking, and storage subsystems and includes developing automation and tools for performance regression tracking and reporting.
Experience Level
Senior - 10+ years of experience in server or system performance architecture or related disciplines.
Responsibilities
Key responsibilities include performance analysis, tooling, and architectural guidance:
- Define and drive server-level performance targets across CPU, GPU, memory, interconnect, networking, and storage.
- Characterize workloads and perform bottleneck analysis using AI training, inference, and HPC benchmarks on NVIDIA and competitive platforms.
- Use profiling, tracing, and analysis tools to root-cause performance issues and identify optimization opportunities.
- Perform trade-off studies on system topology, thermal/power envelopes, and memory hierarchy to guide decisions.
- Collaborate with silicon, platform, firmware, and software teams to identify and close performance gaps through production.
- Develop automation and tooling for performance regression tracking and reporting.
- Represent performance considerations in architecture reviews and cross-functional design discussions.
- Build and maintain analytical performance models and simulation frameworks for next-generation server platforms.
- Publish internal performance studies and best-practice guides for partner and customer enablement.
Requirements
Core technical skills and experience required; additional desirable qualifications follow.
Must-have
- 10+ years of experience in server/system performance architecture or related disciplines.
- Deep understanding of modern server architectures: CPU microarchitecture, PCIe/CXL, DDR/HBM memory subsystems, and coherency protocols.
- Hands-on experience with system-level profiling and performance analysis tools on server platforms.
- Knowledge of GPU-accelerated compute, high-performance networking, or high-performance storage subsystems.
- Proficiency in Python, C/C++, or similar for scripting, data analysis, and tool development.
- Strong communication skills to distill complex performance data into actionable architectural recommendations.
- Proficient using AI-assisted coding and productivity tools to accelerate analysis, automation, and documentation.
Nice-to-have
- Experience with AI/ML training and inference workloads at data-center scale.
- Familiarity with NVIDIA GPU architectures (Hopper, Blackwell, Rubin) and associated software stacks such as CUDA and NCCL.
- Background in chip-to-chip interconnect performance analysis (C2C, UCIe).
- Exposure to power/thermal-aware performance optimization techniques.
- Track record of published performance studies or conference contributions.
Education Requirements
BS, MS, or PhD in Electrical/Computer Engineering, Computer Science, or a related technical field - or equivalent practical experience.
About the Company
Company: NVIDIA
Headquarters: Santa Clara, California, USA
NVIDIA is a global leader in accelerated computing, renowned for its innovative solutions in AI and digital twins that transform diverse industries. The company specializes in networking technologies, providing end-to-end InfiniBand and Ethernet solutions for servers and storage that optimize performance and scalability. NVIDIA serves sectors such as high-performance computing, enterprise data centers, and cloud computing, constantly reinventing its products and services to stay ahead in the market.
