Deep Learning Performance Architect
Work on deep learning performance modeling, analysis, and optimization for inference products focused on large language models (LLMs) and related AI workloads. Influence hardware and software design by identifying performance opportunities and guiding architecture and software teams.
Collaborate with architecture, software, and product teams to specify configurations and metrics that balance performance, power, and accuracy.
Mid-level; requires 3+ years of relevant industry experience.
Key responsibilities include analyzing workloads, building models, and guiding HW/SW design for performance and efficiency.
Must-have technical skills and experience. Education requirements are listed separately below.
BS, MS, or PhD in Computer Science, Electrical Engineering, Mathematics, or a related discipline β or equivalent practical experience.
Company: NVIDIA
Headquarters: Santa Clara, California, USA
NVIDIA is a global leader in accelerated computing, renowned for its innovative solutions in AI and digital twins that transform diverse industries. The company specializes in networking technologies, providing end-to-end InfiniBand and Ethernet solutions for servers and storage that optimize performance and scalability. NVIDIA serves sectors such as high-performance computing, enterprise data centers, and cloud computing, constantly reinventing its products and services to stay ahead in the market.
