Full Stack LLM Engineer
CerebrasJob Title
Full Stack LLM Engineer
Role Summary
Join the Inference Core Model Bringup team to rapidly bring up state-of-the-art open-source and customer models on Cerebras CSX systems. Work across the software stack to translate model architectures, optimize compilers and runtimes, and tune for high-performance inference.
This role focuses on system-level debugging, performance optimization, and integrating models end-to-end for production deployment.
Experience Level
Mid-level — no specific years required; expects demonstrated experience with ML model bringup, compiler/toolchain work, and low-level optimization.
Responsibilities
Primary responsibilities include:
- Bring up ML models end-to-end on Cerebras CSX systems (open-source and customer models).
- Translate model architectures and lower graphs into target compiler IRs and runtimes.
- Implement compiler optimizations and runtime integrations for performance and correctness.
- Profile, debug, and resolve performance, numerical accuracy, and runtime issues spanning model code, compiler IR, runtime behavior, and hardware utilization.
- Prototype and propose tooling, API, or automation improvements to accelerate future bring-ups.
Requirements
Must-have technical skills and experience:
- Proficiency with Python model code and deep learning frameworks such as PyTorch or TensorFlow; familiarity with model internals (attention, MoE, diffusion).
- Proven experience in compiler development and optimizations, especially with LLVM and/or MLIR.
- Strong C/C++ skills and experience with low-level optimization and runtime integration.
- Strong debugging and performance-profiling skills across numerical accuracy, runtime behavior, and hardware utilization.
- Background in optimization techniques for hard combinatorial problems.
- Comfort navigating the full AI toolchain: model code, compiler IRs, profilers, and runtime traces.
Nice-to-have:
- Experience with large-model bringup, model parallelism, and production inference at scale.
Education Requirements
Bachelor’s, Master’s, or PhD in Computer Science, Engineering, or a related field.
About the Company
Company: Cerebras
Headquarters: Sunnyvale, CA, USA
Developer of wafer-scale AI accelerators, Cerebras designs the Wafer Scale Engine (WSE)—one of the world’s largest AI chips—to deliver high-speed training and inference solutions for model labs, enterprises, and AI-native startups.
