Deep Learning Performance Architect
Work on performance modeling, analysis, and optimization for deep learning workloads (including large language models) on state-of-the-art NVIDIA hardware. The role supports processor and system architecture decisions by providing analytical models, prototypes, and performance insight.
Collaborate with architecture, software, and product teams to influence next-generation inference products and improve performance, power, and efficiency.
Mid-level; requires 5+ years of relevant industry experience.
Primary responsibilities focus on analyzing models and defining hardware/software tradeoffs to improve DL inference performance.
Must-haves focused on technical experience and relevant tools.
BS, MS, or PhD in Computer Science, Electrical Engineering, Mathematics, or a related technical field β or equivalent practical experience.
Company: NVIDIA
Headquarters: Santa Clara, California, USA
NVIDIA is a global leader in accelerated computing, renowned for its innovative solutions in AI and digital twins that transform diverse industries. The company specializes in networking technologies, providing end-to-end InfiniBand and Ethernet solutions for servers and storage that optimize performance and scalability. NVIDIA serves sectors such as high-performance computing, enterprise data centers, and cloud computing, constantly reinventing its products and services to stay ahead in the market.
