ML Runtime and Kernel Engineer - Core ML
CerebrasJob Title
ML Runtime and Kernel Engineer - Core ML
Role Summary
The Core ML team builds machine learning algorithms and system support to exploit the Cerebras Wafer-Scale Engine for large-scale training and low-latency inference.
This role implements and optimizes runtime components, distributed execution, and low-level kernels to convert research prototypes into high-performance, production-ready software on Cerebras systems.
Experience Level
Mid-level. The posting does not specify years of experience.
Responsibilities
Work across ML frameworks, compilers, runtimes, and kernel layers to deliver performant end-to-end capabilities for novel ML algorithms.
- Design and implement runtime components and high-performance kernels for Core ML algorithms.
- Translate research prototypes into efficient Cerebras platform implementations and comparative GPU references where useful.
- Profile and debug performance across framework, compiler, runtime, communication, and kernel layers.
- Optimize computation, memory movement, communication, and concurrency for large-scale training and low-latency inference.
- Develop benchmarks, instrumentation, and automated tests for functionality, performance, and numerical correctness.
- Collaborate with researchers and engineering teams to evaluate designs and deliver end-to-end solutions.
- Identify recurring limitations and propose high-leverage platform improvements.
Requirements
Core technical must-haves and desirable skills.
Must-have
- Experience developing high-performance systems software, ML systems, runtimes, compilers, or computational kernels.
- Strong programming skills in C++ and Python.
- Solid understanding of parallel programming, memory management, concurrency, data structures, and performance optimization.
- Proven ability to debug and profile complex software across multiple system layers.
- Familiarity with modern ML frameworks such as PyTorch or JAX and ability to translate algorithmic requirements into reliable software.
Nice-to-have
- Experience with CUDA, Triton, low-level assembly, accelerator programming, or C-like DSLs.
- Experience with compiler internals, distributed runtimes, custom hardware interfaces, or HPC systems.
- Familiarity with LLM training/inference topics: attention, KV-cache management, parallel generation, or distributed execution.
- Experience developing software in research or industrial environments where requirements evolve through experimentation.
- Contributions to significant open-source systems, ML frameworks, compilers, or kernel libraries.
Education Requirements
Bachelor's, Master's, or PhD in Computer Science, Computer Engineering, Electrical Engineering, or a related field, or equivalent practical experience.
About the Company
Company: Cerebras
Headquarters: Sunnyvale, CA, USA
Developer of wafer-scale AI accelerators, Cerebras designs the Wafer Scale Engine (WSE)-one of the world’s largest AI chips-to deliver high-speed training and inference solutions for model labs, enterprises, and AI-native startups.
