Skip to main content
Advanced Micro Devices logo

AI/ML Compiler Developer (NPU Acceleration)

Advanced Micro Devices
September 02, 2026
Full-time
On-site
Hyderabad, Telangana, India
Other Semiconductor Jobs, Level - Senior

Job Title

AI/ML Compiler Developer (NPU Acceleration)

Role Summary

Develop and optimize AI/ML C/C++ kernels and dataflow schedules for AMD Ryzen processors using XDNA Neural Processor Units (NPU). The role focuses on mapping LLMs and Stable Diffusion networks to NPU hardware, improving execution performance, and integrating kernels into the software stack.

Experience Level

Senior-level (Lead / Staff). Experience guidance: approximately 10 years (BS), 8 years (MS), or 5 years (PhD) of relevant engineering experience.

Responsibilities

Key responsibilities include kernel development, performance optimization, validation, and cross-functional collaboration.

  • Design and implement highly optimized C++ kernel libraries for NPU/GPU targets.
  • Map ML workloads (LLMs, Stable Diffusion) to NPU dataflows and schedules.
  • Collaborate with hardware engineers to exploit VLIW/vector core features (MAC, GeMM, non-linear functions).
  • Develop vectorized code using SIMD and ILP techniques for maximum throughput.
  • Profile, analyze, and tune kernel performance; identify and resolve bottlenecks.
  • Develop CPU reference models for ML operators in C++/Python to validate accuracy.
  • Write unit and integration tests and validate kernels across hardware platforms and emulation.
  • Document design specifications, performance improvements, and maintain code via git and PRs.
  • Work with machine learning researchers and software engineers to integrate kernels into the stack.

Requirements

Must-have technical skills and experience for this role, followed by desirable qualifications.

Must-have:

  • Excellent proficiency in C/C++ and Python.
  • Strong understanding of SIMD, tensor/vector, and VLIW processor architectures and how to exploit parallelism.
  • Experience with vectorized programming and parallel computing techniques.
  • Proven experience in performance profiling and low-level optimization of compute kernels.
  • Experience writing unit and integration tests and validating numerical correctness.
  • Good problem-solving skills and focus on performance optimization.

Nice-to-have:

  • Familiarity with ML frameworks such as TensorFlow or PyTorch.
  • Experience with silicon bring-up, pre-silicon validation, or emulation platforms.
  • Knowledge of low-level hardware details (cache hierarchy, memory access patterns).

Education Requirements

BS, MS, or PhD in Computer Science, Electrical Engineering, or a related technical field. The posting specifies approximate experience expectations tied to degree: ~10 years for BS, ~8 years for MS, ~5 years for PhD.


About the Company

Company: Advanced Micro Devices

Headquarters: Sunnyvale, California, USA

Advanced Micro Devices, or AMD, is a global semiconductor company that designs and manufactures microprocessors, graphics processors, and related technologies for a variety of computing devices. Known for pushing the boundaries of innovation, AMD's mission is to deliver high-performance computing solutions for AI, data centers, gaming, and embedded applications. They foster a collaborative, inclusive culture focused on creativity and problem-solving, aiming to drive progress and excellence in technology.

Advanced Micro Devices logo

Date Posted: 2026-09-01