ML Systems Performance Engineer
CerebrasJob Title
ML Systems Performance Engineer
Role Summary
Work on the inference performance team to improve end-to-end model inference speed, throughput, and compute utilization for Cerebras wafer-scale systems. The role spans low-level kernel and compiler optimization, system and cluster runtime performance analysis, performance modeling, and building tooling for diagnostics and visualization.
Experience Level
Mid-level — the posting requests 3+ years of relevant experience in computer architecture, CPU/GPU performance, kernel optimization, or HPC.
Responsibilities
Key responsibilities include:
- Develop kernel-level and end-to-end performance models to estimate performance of state-of-the-art and customer ML models.
- Optimize and debug kernel microcode and compiler algorithms to improve inference speed, throughput, and utilization on the Wafer Scale Engine.
- Analyze and debug runtime performance across system and cluster deployments.
- Design and implement tools and infrastructure to collect, visualize, and analyze performance data from the hardware and compute cluster.
Requirements
Must-have technical skills and experience:
- Strong background in computer architecture.
- Familiarity with low-level deep learning / LLM math.
- 3+ years of experience in relevant domains (computer architecture, CPU/GPU performance, kernel optimization, HPC).
- Experience with CPU/GPU simulators.
- Experience with performance profiling and debugging across system pipelines.
- Proficiency with C++ and Python.
- Strong analytical and problem-solving skills.
Education Requirements
Bachelor's, Master's, or PhD in Electrical Engineering or Computer Science, or equivalent practical experience.
About the Company
Company: Cerebras
Headquarters: Sunnyvale, CA, USA
Developer of wafer-scale AI accelerators, Cerebras designs the Wafer Scale Engine (WSE)—one of the world’s largest AI chips—to deliver high-speed training and inference solutions for model labs, enterprises, and AI-native startups.
