Skip to main content
Cerebras logo

Director/Sr. Manager, AI Inference Model Scaling

Cerebras
August 27, 2026
Full-time
Remote friendly (Sunnyvale, California, United States)
Worldwide
EDA Jobs, Level - Senior

Job Title

Director / Senior Manager, AI Inference Model Scaling

Role Summary

Lead and scale the Inference Model Scaling organization responsible for enabling foundation and generative AI models on Cerebras' Wafer-Scale Engine (WSE). Define technical vision, organizational strategy, and execution plans for model compilation, graph optimization, high-performance kernels, and runtime integration.

Work across compiler, runtime, cloud, hardware, product, and research teams to bring new model architectures into production and improve inference performance at scale.

Experience Level

Senior — typically 12+ years building compiler, ML systems, or infrastructure software and 5+ years leading engineering teams.

Responsibilities

Drive technical and organizational execution to enable state-of-the-art models on Cerebras hardware.

  • Define and communicate the team technical roadmap, strategy, and engineering standards.
  • Set technical direction across multiple teams and lead design reviews.
  • Prioritize and deliver support for emerging LLM architectures and inference workloads.
  • Hire, mentor, and grow a distributed engineering organization; develop technical leaders and managers.
  • Plan headcount, drive organizational priorities, and scale processes while maintaining velocity.
  • Collaborate with Cloud Platform, ML, Hardware, Product, and customer-facing teams for end-to-end enablement.
  • Influence hardware/software co-design through model enablement and optimization insights.
  • Own planning, prioritization, and predictable delivery across multiple concurrent initiatives.

Requirements

Key must-have skills and experience for immediate impact, followed by desirable qualifications.

  • Must-have: 12+ years building compilers, ML systems, or infrastructure software.
  • 5+ years managing and scaling engineering teams.
  • Deep experience with modern compiler infrastructure (LLVM, MLIR, XLA, TVM, Torch FX, or similar).
  • Strong understanding of graph compilation and optimization techniques.
  • Proficient in Python and C++ and experienced delivering production-quality software.
  • Strong communication and cross-functional leadership skills.

Nice-to-have:

  • Experience building compiler frontends or optimizations for AI accelerators.
  • Familiarity with PyTorch, JAX, TensorFlow, or ONNX and LLM inference/training systems.
  • Experience with distributed compilation and working closely with hardware architects.
  • Proven track record leading teams through rapid growth.

Education Requirements

BS, MS, or PhD in Computer Science, Computer Engineering, or a related technical field.


About the Company

Company: Cerebras

Headquarters: Sunnyvale, CA, USA

Developer of wafer-scale AI accelerators, Cerebras designs the Wafer Scale Engine (WSE)—one of the world’s largest AI chips—to deliver high-speed training and inference solutions for model labs, enterprises, and AI-native startups.

Cerebras logo

Date Posted: 2026-08-27