Skip to main content
Cerebras logo

ML Runtime and Kernel Engineer - Core ML

Cerebras
September 29, 2026
Full-time
On-site
Sunnyvale, California, United States
EDA Jobs, Level - Mid-Career

Job Title

ML Runtime and Kernel Engineer - Core ML

Role Summary

The Core ML team builds machine learning algorithms and system support to exploit the Cerebras Wafer-Scale Engine for large-scale training and low-latency inference.

This role implements and optimizes runtime components, distributed execution, and low-level kernels to convert research prototypes into high-performance, production-ready software on Cerebras systems.

Experience Level

Mid-level. The posting does not specify years of experience.

Responsibilities

Work across ML frameworks, compilers, runtimes, and kernel layers to deliver performant end-to-end capabilities for novel ML algorithms.

  • Design and implement runtime components and high-performance kernels for Core ML algorithms.
  • Translate research prototypes into efficient Cerebras platform implementations and comparative GPU references where useful.
  • Profile and debug performance across framework, compiler, runtime, communication, and kernel layers.
  • Optimize computation, memory movement, communication, and concurrency for large-scale training and low-latency inference.
  • Develop benchmarks, instrumentation, and automated tests for functionality, performance, and numerical correctness.
  • Collaborate with researchers and engineering teams to evaluate designs and deliver end-to-end solutions.
  • Identify recurring limitations and propose high-leverage platform improvements.

Requirements

Core technical must-haves and desirable skills.

Must-have

  • Experience developing high-performance systems software, ML systems, runtimes, compilers, or computational kernels.
  • Strong programming skills in C++ and Python.
  • Solid understanding of parallel programming, memory management, concurrency, data structures, and performance optimization.
  • Proven ability to debug and profile complex software across multiple system layers.
  • Familiarity with modern ML frameworks such as PyTorch or JAX and ability to translate algorithmic requirements into reliable software.

Nice-to-have

  • Experience with CUDA, Triton, low-level assembly, accelerator programming, or C-like DSLs.
  • Experience with compiler internals, distributed runtimes, custom hardware interfaces, or HPC systems.
  • Familiarity with LLM training/inference topics: attention, KV-cache management, parallel generation, or distributed execution.
  • Experience developing software in research or industrial environments where requirements evolve through experimentation.
  • Contributions to significant open-source systems, ML frameworks, compilers, or kernel libraries.

Education Requirements

Bachelor's, Master's, or PhD in Computer Science, Computer Engineering, Electrical Engineering, or a related field, or equivalent practical experience.


About the Company

Company: Cerebras

Headquarters: Sunnyvale, CA, USA

Developer of wafer-scale AI accelerators, Cerebras designs the Wafer Scale Engine (WSE)-one of the world’s largest AI chips-to deliver high-speed training and inference solutions for model labs, enterprises, and AI-native startups.

Cerebras logo

Date Posted: 2026-09-28