ML Systems Integration Engineer
CerebrasJob Title
ML Systems Integration Engineer
Role Summary
The ML Systems Integration Engineer will support bring-up, validation, and debugging of next-generation AI hardware systems and their supporting software. The role focuses on reproducing and diagnosing system-level failures, building automation and tooling to improve validation workflows, and collaborating with hardware and firmware teams to drive systems toward production readiness.
Experience Level
Mid-level. No explicit years of experience specified.
Responsibilities
Primary responsibilities include system bring-up, debugging, and automation to accelerate hardware validation and deployment.
- Participate in bring-up of new AI hardware systems and associated software infrastructure.
- Reproduce, triage, and diagnose complex system-level issues spanning hardware and software.
- Investigate failures using logs, telemetry, and diagnostic tools to identify root causes.
- Develop software and test frameworks to validate and stress distributed hardware systems.
- Build automation and internal tooling to improve validation, observability, and debugging workflows.
- Collaborate closely with hardware, firmware, and systems engineers to isolate integration issues.
- Support validation and qualification of new hardware generations toward production readiness.
- Continuously improve engineering workflows related to debugging, testing, and automation.
Requirements
Must-have technical skills and behaviors for the role.
- Strong programming skills in Python and/or C++.
- Excellent debugging and systematic problem-solving ability for complex technical issues.
- Solid understanding of operating systems fundamentals: processes, threads, memory management, concurrency, and IPC.
- Experience working in Linux development environments.
- Understanding of computer architecture and hardware–software interactions.
- Ability to work effectively across multiple engineering teams and communicate technical issues clearly.
Nice-to-have:
- Experience building automation frameworks, internal tooling, or test infrastructure.
- Familiarity with distributed systems concepts and networking fundamentals.
- Experience debugging large-scale systems, performance analysis, system telemetry, or log analysis.
- Exposure to production systems validation or infrastructure reliability engineering.
Education Requirements
BS or MS in Computer Science, Computer Engineering, Electrical Engineering, or a related technical field, as stated in the posting.
About the Company
Company: Cerebras
Headquarters: Sunnyvale, CA, USA
Developer of wafer-scale AI accelerators, Cerebras designs the Wafer Scale Engine (WSE)—one of the world’s largest AI chips—to deliver high-speed training and inference solutions for model labs, enterprises, and AI-native startups.
