Senior GPU Performance Software Engineer
Intel CorporationJob Title
Senior GPU Performance Software Engineer
Role Summary
Develop and optimize low-level GPU kernels, math primitives, and runtime infrastructure for oneDNN to accelerate AI frameworks on Intel GPUs. Collaborate with hardware, compiler, and framework teams to shape kernel architectures and improve performance at scale.
This is a low-level software engineering and hardware-acceleration role; it does not involve building or training ML models.
Experience Level
Senior-level. Requires multi-year professional software development experience; see Requirements for specific years and domain guidance.
Responsibilities
Deliver high-performance GPU primitives and the supporting tooling and infrastructure.
- Develop high-performance GEMM, convolution, and attention kernels and scalable JIT/codegen infrastructure for GPU kernel generation.
- Implement fusion, memory-traffic, and mixed-precision/quantized execution optimizations (e.g., BF16, FP16, INT8, FP8, FP4).
- Build analytical and empirical performance models; profile and eliminate bottlenecks across oneDNN GPU primitives and runtime paths.
- Co-design GPU primitives and kernel architectures with hardware and compiler teams to influence next-generation accelerator capabilities.
- Improve validation, benchmarking, and CI infrastructure for performance-critical GPU workloads.
Requirements
Must-have technical skills and experience.
- Expert-level modern C++ with 5+ years of professional software development experience.
- At least 2+ years hands-on programming and kernel optimization on GPUs (SYCL/DPC++, OpenCL, CUDA, or HIP) OR 5+ years of low-level performance optimization experience on CPUs.
- Strong foundations in computer architecture, cache hierarchies, memory subsystems, and parallel programming paradigms (multi-threading, SIMD/vectorization).
- Experience with performance modeling, profiling, and eliminating runtime bottlenecks.
Nice-to-have:
- Experience developing high-performance math libraries (GEMM, convolution, reduction, FFT).
- GPU assembly-level tuning or compiler optimization experience.
- Familiarity with OpenMP, oneTBB, or other parallel programming APIs.
- Basic understanding of deep-learning primitives and their use in upstream frameworks.
Education Requirements
BSc, MSc, or PhD in Computer Science, Computer Engineering, Mathematics, Physics, or a highly technical related field.
About the Company
Company: Intel Corporation
Headquarters: Santa Clara, California, USA
Intel Corporation is a leading multinational technology company known for its innovative semiconductor solutions, including microprocessors, artificial intelligence accelerators, and memory products. Headquartered in the United States, Intel focuses on cutting-edge technology and a collaborative working environment, driving advancements in semiconductor manufacturing to meet global demands. The company emphasizes professional development and aims to shape the future of technology through groundbreaking designs.
