Director/Sr. Manager, AI Inference Model Scaling
CerebrasJob Title
Director / Senior Manager, AI Inference Model Scaling
Role Summary
Lead and scale the Inference Model Scaling organization responsible for enabling foundation and generative AI models on Cerebras' Wafer-Scale Engine (WSE). Define technical vision, organizational strategy, and execution plans for model compilation, graph optimization, high-performance kernels, and runtime integration.
Work across compiler, runtime, cloud, hardware, product, and research teams to bring new model architectures into production and improve inference performance at scale.
Experience Level
Senior — typically 12+ years building compiler, ML systems, or infrastructure software and 5+ years leading engineering teams.
Responsibilities
Drive technical and organizational execution to enable state-of-the-art models on Cerebras hardware.
- Define and communicate the team technical roadmap, strategy, and engineering standards.
- Set technical direction across multiple teams and lead design reviews.
- Prioritize and deliver support for emerging LLM architectures and inference workloads.
- Hire, mentor, and grow a distributed engineering organization; develop technical leaders and managers.
- Plan headcount, drive organizational priorities, and scale processes while maintaining velocity.
- Collaborate with Cloud Platform, ML, Hardware, Product, and customer-facing teams for end-to-end enablement.
- Influence hardware/software co-design through model enablement and optimization insights.
- Own planning, prioritization, and predictable delivery across multiple concurrent initiatives.
Requirements
Key must-have skills and experience for immediate impact, followed by desirable qualifications.
- Must-have: 12+ years building compilers, ML systems, or infrastructure software.
- 5+ years managing and scaling engineering teams.
- Deep experience with modern compiler infrastructure (LLVM, MLIR, XLA, TVM, Torch FX, or similar).
- Strong understanding of graph compilation and optimization techniques.
- Proficient in Python and C++ and experienced delivering production-quality software.
- Strong communication and cross-functional leadership skills.
Nice-to-have:
- Experience building compiler frontends or optimizations for AI accelerators.
- Familiarity with PyTorch, JAX, TensorFlow, or ONNX and LLM inference/training systems.
- Experience with distributed compilation and working closely with hardware architects.
- Proven track record leading teams through rapid growth.
Education Requirements
BS, MS, or PhD in Computer Science, Computer Engineering, or a related technical field.
About the Company
Company: Cerebras
Headquarters: Sunnyvale, CA, USA
Developer of wafer-scale AI accelerators, Cerebras designs the Wafer Scale Engine (WSE)—one of the world’s largest AI chips—to deliver high-speed training and inference solutions for model labs, enterprises, and AI-native startups.
