Skip to main content
Cerebras logo

ML Systems Performance Engineer

Cerebras
August 27, 2026
Full-time
On-site
Sunnyvale, California, United States
SoC Architecture Jobs, Level - Mid-Career

Job Title

ML Systems Performance Engineer

Role Summary

Work on the inference performance team to improve end-to-end model inference speed, throughput, and compute utilization for Cerebras wafer-scale systems. The role spans low-level kernel and compiler optimization, system and cluster runtime performance analysis, performance modeling, and building tooling for diagnostics and visualization.

Experience Level

Mid-level — the posting requests 3+ years of relevant experience in computer architecture, CPU/GPU performance, kernel optimization, or HPC.

Responsibilities

Key responsibilities include:

  • Develop kernel-level and end-to-end performance models to estimate performance of state-of-the-art and customer ML models.
  • Optimize and debug kernel microcode and compiler algorithms to improve inference speed, throughput, and utilization on the Wafer Scale Engine.
  • Analyze and debug runtime performance across system and cluster deployments.
  • Design and implement tools and infrastructure to collect, visualize, and analyze performance data from the hardware and compute cluster.

Requirements

Must-have technical skills and experience:

  • Strong background in computer architecture.
  • Familiarity with low-level deep learning / LLM math.
  • 3+ years of experience in relevant domains (computer architecture, CPU/GPU performance, kernel optimization, HPC).
  • Experience with CPU/GPU simulators.
  • Experience with performance profiling and debugging across system pipelines.
  • Proficiency with C++ and Python.
  • Strong analytical and problem-solving skills.

Education Requirements

Bachelor's, Master's, or PhD in Electrical Engineering or Computer Science, or equivalent practical experience.


About the Company

Company: Cerebras

Headquarters: Sunnyvale, CA, USA

Developer of wafer-scale AI accelerators, Cerebras designs the Wafer Scale Engine (WSE)—one of the world’s largest AI chips—to deliver high-speed training and inference solutions for model labs, enterprises, and AI-native startups.

Cerebras logo

Date Posted: 2026-08-27