Meta logo

Software Engineer, AI Kernels & Performance Optimization — MTIA Software

Meta
August 14, 2026
Full-time
On-site
Bellevue, Washington, United States
$183,997 - $257,000 USD yearly
EDA Jobs, Level - Senior

Job Title

Software Engineer, AI Kernels & Performance Optimization — MTIA Software

Role Summary

Work on AI kernel and optimization software for Meta's MTIA accelerators, owning performance from architecture analysis through production deployment. The team builds kernel libraries, authoring frameworks, compiler interfaces, and PyTorch integration to deliver high-performance kernels for recommendation, ranking, and generative AI workloads.

Experience Level

Senior-level role; the posting targets experienced engineers (typical guidance: 6+ years; senior-level responsibilities and technical leadership expected).

Responsibilities

Hands-on engineering role focused on kernel performance, numerics, and hardware/software co-design.

  • Design, implement, and optimize compute and communication kernels (GEMM, attention variants, normalization, fused ops, sparse/quantized paths) for MTIA accelerators.
  • Profile and root-cause performance across the stack (instruction scheduling, memory hierarchy, DMA, on-chip interconnect, collectives) and drive fixes at the correct layer.
  • Build and extend kernel authoring frameworks, templates, and libraries to broaden high-performance coverage.
  • Deliver and maintain PyTorch operator coverage across eager and compiled execution paths for production models.
  • Collaborate with silicon architecture and design teams to quantify feature value, characterize rooflines pre-silicon, and influence hardware decisions.
  • Develop software mitigations for hardware limitations and generalize solutions for the organization.
  • Mentor engineers and set technical direction for a kernel domain; write design documents and align cross-functional partners.

Requirements

Must-have technical skills and experience for successful performance engineering on accelerators.

  • 6+ years professional experience in high-performance computing, accelerator kernel development, compiler backends, or systems performance engineering.
  • Proficiency in C++ and Python, including low-level systems programming, templates/generic programming, and performance-critical code.
  • Demonstrated experience writing and optimizing kernels for parallel architectures (GPU CUDA/ROCm/HIP, TPU/AI ASICs, or SIMD/vector CPU targets).
  • Working knowledge of computer architecture relevant to performance: memory hierarchies, bandwidth/latency tradeoffs, occupancy/scheduling, vectorization, and synchronization.
  • Measurement-driven performance methodology: build roofline or analytical models, profile against them, and explain residual gaps.
  • Ability to read hardware specifications and RTL-adjacent documentation and to identify whether issues are hardware- or software-rooted.
  • Experience with numerics and low-precision formats (FP8, block-scaled types, integer quantization) is highly desirable.
  • Familiarity with compiler/codegen technologies (MLIR, LLVM, TVM, XLA, Halide) and with distributed execution/collectives is a plus.

Education Requirements

Bachelor's degree in Computer Science, Computer Engineering, or a related technical field, or equivalent practical experience. Advanced degree (master's/PhD) or equivalent experience is noted as preferred in the posting.


About the Company

Company: Meta

Headquarters: Menlo Park, California, United States

Meta is a technology company focused on connecting people and building immersive experiences through virtual and augmented reality. It develops products and services that encompass social networking, communication, and advanced display technologies. As a leader in the tech industry, Meta continually innovates to enhance user interaction in the digital realm.

Meta logo

Date Posted: 2026-08-13