AI Inference Core - SDET Technical Lead, Release Integration Testing
CerebrasJob Title
AI Inference Core - SDET Technical Lead, Release Integration Testing
Role Summary
Technical lead responsible for establishing and operating Release Integration Testing (RIT) for the AI Inference Core. Define and enforce integration and release readiness criteria across software, infrastructure, and hardware layers to ensure reliable production releases.
Work hands-on across AI frameworks, runtimes, compilers, kernels, distributed systems, infrastructure, and hardware. Role requires in-office presence at least three days per week; offices in Sunnyvale, CA and Toronto, ON.
Experience Level
Senior-level technical lead. The posting identifies a technical leadership role (title: Technical Lead); specific years of experience are not specified.
Responsibilities
Lead the design and operation of release integration testing and evidence-based release gating for inference software and hardware.
- Define RIT strategy, engagement criteria, ownership boundaries, entry/exit criteria, and coverage expectations for AI Inference Core.
- Engage early on high-risk inference changes and identify cross-stack dependencies and interaction risks.
- Own the inference-path readiness gate by reviewing unit, simulation, benchmark, feature-test, and integration evidence and documenting gaps.
- Lead integrated end-to-end inference validation across the cloud-to-wafer stack and add risk-based scenarios to release regression.
- Improve master and release-branch stability through metrics, failure classification, reporting, dashboards, qualification workflows, and pipelines.
- Lead first-pass regression and rollout triage; coordinate owners, drive RCA, and assign missing coverage to the correct layer.
- Mentor and partner with SDETs, feature teams, integration, core infra, release owners, and deployment teams.
- Advance automation, diagnostics, probes, and test-roadmap planning between active engagements.
Requirements
Clear distinction between required capabilities and preferred experience.
Must-have- Strong software-engineering fundamentals and programming ability in Python, Go, or a similar language.
- Demonstrated technical leadership in software quality, test infrastructure, systems validation, release engineering, or complex software integration.
- Experience designing automation and test architecture for distributed, systems-level, infrastructure, or AI software.
- Proven ability to break down ambiguous cross-stack failures, form hypotheses, gather evidence, and drive issues to resolution.
- Strong understanding of risk-based testing, release readiness, regression strategy, failure analysis, and quality metrics.
- Ability to influence and align multiple engineering teams and communicate technical risk clearly under pressure.
- Experience with software/hardware co-design, hardware accelerators, compilers, kernels, runtimes, or low-level systems.
- Experience with AI infrastructure, model deployment, LLMs, multimodal workloads, or large-scale compute clusters.
- Experience building distributed test frameworks, release pipelines, dashboards, performance testing, observability, fault injection, or production failure analysis.
- Familiarity with containers, cluster orchestration, cloud infrastructure, CI/CD, or high-performance computing.
- Track record of building and scaling a release or quality capability in a fast-moving environment.
Education Requirements
Not specified.
About the Company
Company: Cerebras
Headquarters: Sunnyvale, CA, USA
Developer of wafer-scale AI accelerators, Cerebras designs the Wafer Scale Engine (WSE)—one of the world’s largest AI chips—to deliver high-speed training and inference solutions for model labs, enterprises, and AI-native startups.
