Skip to main content
Cerebras logo

Hardware Analytics Engineer

Cerebras
August 27, 2026
Full-time
On-site
Sunnyvale, California, United States
$213,675 - $225,000 USD yearly
Test Engineering Jobs, Level - Mid-Career

Job Title

Hardware Analytics Engineer

Role Summary

Design and operate large-scale data and analytics systems to monitor, analyze, and improve hardware performance, reliability, and efficiency for AI server platforms. Work with hardware, firmware, and datacenter teams to turn multi‑terabyte telemetry into actionable diagnostics, anomaly detection, and prescriptive recommendations.

This role focuses on telemetry pipelines, ML-driven failure prediction, and experiments to optimize power, thermal, and reliability for next‑generation AI infrastructure.

Experience Level

Mid-level — requires a minimum of 3 years of relevant professional experience.

Responsibilities

Deliver scalable telemetry ingestion, analysis, and visualization systems and lead analytics to improve hardware reliability and performance.

  • Design and optimize scalable data pipeline architectures and ETL for multi‑terabyte hardware telemetry.
  • Build and maintain distributed data processing frameworks (Hive, Spark) to aggregate and analyze utilization, power, thermal, acoustic, and reliability metrics.
  • Develop ML models and statistical methods for anomaly detection, failure prediction, and prescriptive optimization.
  • Engineer real‑time telemetry ingestion, monitoring, and dashboards to surface hardware health to engineering and operations teams.
  • Lead hardware characterization and A/B studies (thermal/cooling) to define operational envelopes and efficiency improvements.
  • Perform root cause analysis on systemic failures and implement fixes across CPU, GPU, DRAM, PCIe, networking, and storage subsystems.
  • Define and operationalize custom efficiency and reliability metrics to improve scalability, energy efficiency, and sustainability.
  • Collaborate cross‑functionally to support evolution of next‑generation AI platforms and silicon products.

Requirements

Core technical skills and experience required for day‑to‑day success. Education degree requirements are listed below in Education Requirements.

  • Must-have: Proven experience with large‑scale data pipeline architecture, ETL, and distributed data processing (Hive, Spark).
  • Must-have: Strong Python and SQL skills; experience with Linux and automation scripting.
  • Must-have: Experience designing, training, and deploying ML models for hardware performance optimization and failure prediction; predictive modeling, anomaly detection, and A/B testing.
  • Must-have: Dashboarding and visualization experience (Tableau or equivalent) to present hardware health and reliability metrics.
  • Must-have: Domain experience in hardware analytics for compute, storage, and AI servers, including power and thermal optimization and reliability modeling for components (CPU, GPU, DRAM, SSD).
  • Nice-to-have: Experience with GPU burn‑in processes, hyperscale telemetry frameworks, and sustainability metrics (energy/water efficiency).
  • Nice-to-have: Prior collaboration with firmware, datacenter operations, or hardware validation teams on troubleshooting and systemic fixes.

Education Requirements

Master's degree (or foreign equivalent) in Electrical Engineering, Computer Engineering, Computer Science, or a related field, plus a minimum of 3 years of relevant experience as a Hardware Analytics Engineer, Hardware Engineer, Data Engineer, or similar role.


About the Company

Company: Cerebras

Headquarters: Sunnyvale, CA, USA

Developer of wafer-scale AI accelerators, Cerebras designs the Wafer Scale Engine (WSE)—one of the world’s largest AI chips—to deliver high-speed training and inference solutions for model labs, enterprises, and AI-native startups.

Cerebras logo

Date Posted: 2026-08-27