ML Software Tool Development Engineer
CerebrasJob Title
ML Software Tool Development Engineer
Role Summary
Design and implement system-level debugging, validation, and observability platforms and tools for machine-learning software running on Cerebras hardware. Collaborate with compiler, hardware, firmware, runtime, and infrastructure teams to improve bring-up, profiling, and incident response.
Work focuses on enabling reliable, high-performance tooling and diagnostics for large-scale ML training and inference on custom AI accelerators.
Experience Level
Mid-level. Years of experience not specified; the role expects independent delivery and leadership of complex technical projects.
Responsibilities
Key responsibilities include implementing tooling and platforms to detect, analyze, and remediate system-level issues.
- Lead design and implementation of system-level debugging, validation, and observability platforms.
- Develop automated systems for collecting and analyzing numerical and execution anomalies.
- Create visualization and analysis tools for efficient root-cause investigation.
- Build frameworks for failure classification, regression detection, and anomaly monitoring.
- Extend compilers, runtimes, and programming interfaces to support profiling and instrumentation.
- Improve system bring-up, low-level debug, and validation workflows.
- Partner cross-functionally with compiler, hardware, firmware, runtime, and infrastructure teams.
- Establish and promote best practices for debuggability, reliability, and operational excellence.
- Lead high-impact initiatives and support incident response with long-term corrective actions.
Requirements
Must-have qualifications are listed first; preferred skills follow.
- Must-have: Strong proficiency in C++ and Python with experience building reliable, high-performance systems and tooling.
- Must-have: Demonstrated experience debugging complex hardware/software systems and driving issues to root cause.
- Must-have: Experience analyzing system-level data structures, execution graphs, or dependency networks for diagnostics and validation.
- Must-have: Proven ability to design and build intuitive visualization and analysis tools for complex technical data.
- Must-have: Experience with compiler internals, custom hardware interfaces, or low-level protocol design.
- Must-have: Strong written and verbal communication skills and ability to lead technical projects end-to-end.
- Preferred: Familiarity with machine learning training and inference pipelines, especially distributed training and large-model scaling.
- Preferred: Prior work on high-performance clusters, HPC systems, or custom hardware/software co-design.
Education Requirements
Not specified.
About the Company
Company: Cerebras
Headquarters: Sunnyvale, CA, USA
Developer of wafer-scale AI accelerators, Cerebras designs the Wafer Scale Engine (WSE)—one of the world’s largest AI chips—to deliver high-speed training and inference solutions for model labs, enterprises, and AI-native startups.
