CoDesign & NextGen Performance Engineer
CerebrasJob Title
CoDesign & NextGen Performance Engineer
Role Summary
This role focuses on characterizing, analyzing, and optimizing the performance of state-of-the-art AI models running on Cerebras hardware. The engineer will work across hardware and software (kernel, compiler, runtime, cluster) to identify bottlenecks, improve computational efficiency, and influence next-generation architecture and software systems.
Experience Level
Mid-level. Requires approximately 3+ years of relevant experience in computer architecture, CPU/GPU performance, kernel optimization, or HPC.
Responsibilities
Primary responsibilities include:
- Bring up and optimize performance on new generations of the Cerebras Wafer Scale Engine (WSE).
- Build kernel-level and end-to-end performance models to estimate model performance.
- Optimize and debug kernel microcode and compiler algorithms to improve inference speed, throughput, and compute utilization.
- Investigate and resolve runtime and cluster-level performance issues.
- Develop tools and infrastructure to collect, analyze, and visualize performance data from the WSE and compute clusters.
Requirements
Must-have skills and experience:
- Strong background in computer architecture.
- Familiarity with low-level deep learning / LLM math.
- 3+ years of relevant experience (architecture, CPU/GPU performance, kernel optimization, or HPC).
- Experience with CPU/GPU simulators.
- Experience with performance profiling and debugging across system pipelines.
- Proficiency in C++ and Python.
- Strong analytical and problem-solving mindset.
Nice-to-have:
- Experience with compiler optimization, kernel microcode development, or cluster-scale performance tooling.
Education Requirements
Bachelor's, Master's, or PhD in Electrical Engineering, Computer Science, or a related technical field (as listed in the posting).
About the Company
Company: Cerebras
Headquarters: Sunnyvale, CA, USA
Developer of wafer-scale AI accelerators, Cerebras designs the Wafer Scale Engine (WSE)—one of the world’s largest AI chips—to deliver high-speed training and inference solutions for model labs, enterprises, and AI-native startups.
