Skip to main content
Cerebras logo

Staff Kernel Optimization Engineer

Cerebras
August 27, 2026
Full-time
Remote
Worldwide
Other Semiconductor Jobs, Level - Senior

Job Title

Staff Kernel Optimization Engineer

Role Summary

Develop and optimize high-performance ML and HPC kernel libraries that fully leverage Cerebras’ massively parallel processor architecture. Work on kernel design, performance tuning, validation, and integration with system-level architecture to maximize compute utilization for training and inference workloads.

Experience Level

Senior / Staff level. The role expects experienced engineers capable of driving low-level kernel design and cross-functional performance work; no specific years-of-experience stated.

Responsibilities

Design, implement, and validate highly optimized kernel routines and algorithms to run on Cerebras hardware.

  • Design specifications and mappings for ML and linear-algebra kernels using parallel programming algorithms.
  • Implement and debug high-performance kernel routines in low-level assembly and a C-like domain-specific language (CSL).
  • Develop and maintain a kernel library of parallel and distributed algorithms to maximize utilization and training efficiency.
  • Use mathematical models and performance analysis to guide design and optimization decisions.
  • Develop unit and system tests to verify correctness and performance of kernel libraries.
  • Collaborate with chip and system architects to optimize instruction sets, microarchitecture, and IO for next-generation systems.
  • Monitor emerging ML application trends and evolve kernel architecture to meet new computational challenges.

Requirements

Key technical skills and capabilities required or strongly preferred.

  • Strong understanding of hardware architecture concepts and ability to learn new architectures quickly. (must-have)
  • Proficient in C++ and Python for systems and tooling development. (must-have)
  • Experience developing libraries or APIs and following best practices for maintainable interfaces. (must-have)
  • Strong debugging skills and experience diagnosing issues across complex software stacks. (must-have)
  • Ability to develop and debug low-level assembly and domain-specific language kernels or willingness to acquire this expertise quickly. (must-have)
  • Preferred: prior kernel development or testing experience.
  • Preferred: familiarity with parallel algorithms and distributed-memory systems.
  • Preferred: experience programming accelerators (GPUs, FPGAs) and optimizing HPC kernels.
  • Preferred: familiarity with machine learning frameworks (e.g., TensorFlow, PyTorch) and neural network workloads.

Education Requirements

Bachelor's, Master's, or PhD (or foreign equivalents) in Computer Science, Computer Engineering, Mathematics, or related technical fields.


About the Company

Company: Cerebras

Headquarters: Sunnyvale, CA, USA

Developer of wafer-scale AI accelerators, Cerebras designs the Wafer Scale Engine (WSE)—one of the world’s largest AI chips—to deliver high-speed training and inference solutions for model labs, enterprises, and AI-native startups.

Cerebras logo

Date Posted: 2026-08-27