AI/ML Compiler Developer (NPU Acceleration)
Advanced Micro DevicesJob Title
AI/ML Compiler Developer (NPU Acceleration)
Role Summary
Develop and optimize AI/ML C/C++ kernels and dataflow schedules for AMD Ryzen processors using XDNA Neural Processor Units (NPU). The role focuses on mapping LLMs and Stable Diffusion networks to NPU hardware, improving execution performance, and integrating kernels into the software stack.
Experience Level
Senior-level (Lead / Staff). Experience guidance: approximately 10 years (BS), 8 years (MS), or 5 years (PhD) of relevant engineering experience.
Responsibilities
Key responsibilities include kernel development, performance optimization, validation, and cross-functional collaboration.
- Design and implement highly optimized C++ kernel libraries for NPU/GPU targets.
- Map ML workloads (LLMs, Stable Diffusion) to NPU dataflows and schedules.
- Collaborate with hardware engineers to exploit VLIW/vector core features (MAC, GeMM, non-linear functions).
- Develop vectorized code using SIMD and ILP techniques for maximum throughput.
- Profile, analyze, and tune kernel performance; identify and resolve bottlenecks.
- Develop CPU reference models for ML operators in C++/Python to validate accuracy.
- Write unit and integration tests and validate kernels across hardware platforms and emulation.
- Document design specifications, performance improvements, and maintain code via git and PRs.
- Work with machine learning researchers and software engineers to integrate kernels into the stack.
Requirements
Must-have technical skills and experience for this role, followed by desirable qualifications.
Must-have:
- Excellent proficiency in C/C++ and Python.
- Strong understanding of SIMD, tensor/vector, and VLIW processor architectures and how to exploit parallelism.
- Experience with vectorized programming and parallel computing techniques.
- Proven experience in performance profiling and low-level optimization of compute kernels.
- Experience writing unit and integration tests and validating numerical correctness.
- Good problem-solving skills and focus on performance optimization.
Nice-to-have:
- Familiarity with ML frameworks such as TensorFlow or PyTorch.
- Experience with silicon bring-up, pre-silicon validation, or emulation platforms.
- Knowledge of low-level hardware details (cache hierarchy, memory access patterns).
Education Requirements
BS, MS, or PhD in Computer Science, Electrical Engineering, or a related technical field. The posting specifies approximate experience expectations tied to degree: ~10 years for BS, ~8 years for MS, ~5 years for PhD.
About the Company
Company: Advanced Micro Devices
Headquarters: Sunnyvale, California, USA
Advanced Micro Devices, or AMD, is a global semiconductor company that designs and manufactures microprocessors, graphics processors, and related technologies for a variety of computing devices. Known for pushing the boundaries of innovation, AMD's mission is to deliver high-performance computing solutions for AI, data centers, gaming, and embedded applications. They foster a collaborative, inclusive culture focused on creativity and problem-solving, aiming to drive progress and excellence in technology.
