Senior Staff AI Accelerator Performance Architect
CerebrasJob Title
Senior Staff AI Accelerator Performance Architect
Role Summary
Lead performance modeling and architecture evaluation for next-generation wafer-scale AI accelerator systems. Connect real workloads to architectural behavior, quantify bottlenecks, and produce recommendations that influence hardware and software roadmaps.
Base salary range: $175,000 to $275,000 annually; actual compensation may include bonus and equity.
Experience Level
Senior-level. Typical background: 7+ years in performance analysis, performance modeling, or architecture exploration for CPUs, GPUs, AI accelerators, or high-performance computing systems.
Responsibilities
Primary responsibilities include building and maintaining performance models, analyzing workloads end-to-end, and translating findings into architectural recommendations.
- Own and evolve performance models and modeling methodologies for accelerator and system architectures.
- Build analytical, simulation-based, or trace-driven models across workloads and product generations.
- Analyze kernels and end-to-end training and inference to locate time, bandwidth, compute and capacity bottlenecks.
- Quantify opportunities to improve latency, throughput, utilization and energy efficiency.
- Evaluate proposed architectural features for expected performance return across representative workloads.
- Study mapping of models and kernels onto compute, memory and communication architectures.
- Partner with architecture, compiler, kernel, runtime and systems teams to evaluate mappings and optimizations.
- Validate and correlate models with RTL, emulation, FPGA prototypes or silicon measurements.
- Create concise, actionable recommendations and workload projections grounded in transparent assumptions.
Requirements
Must-have technical skills and experience relevant to accelerator performance and system architecture.
- 7+ years of experience in performance analysis, modeling, or architecture exploration for CPUs, GPUs, AI accelerators or HPC systems.
- Strong hardware-architecture knowledge acquired via hardware, compiler, kernel, runtime or system-performance work.
- Experience developing analytical, simulation-based or trace-driven performance models using Python, C++ or similar.
- Solid understanding of processor architecture, memory systems, interconnects, parallel execution and hardware resource constraints.
- Ability to move between kernel-level behavior and end-to-end application or system performance; experience profiling and validating hypotheses with quantitative evidence.
- Experience with kernel optimization, compiler performance, runtime scheduling, or distributed accelerator systems is highly relevant.
- Clear communication of modeling assumptions, uncertainty, bottlenecks and recommendations.
- Comfort reasoning about microarchitecture and collaborating with RTL and physical-design teams (not responsible for production RTL ownership).
Education Requirements
MS or PhD in Electrical Engineering, Computer Engineering, Computer Science, or equivalent practical experience.
About the Company
Company: Cerebras
Headquarters: Sunnyvale, CA, USA
Developer of wafer-scale AI accelerators, Cerebras designs the Wafer Scale Engine (WSE)-one of the world’s largest AI chips-to deliver high-speed training and inference solutions for model labs, enterprises, and AI-native startups.
