Infrastructure Software Engineer
EtchedJob Title
Infrastructure Software Engineer
Role Summary
Build and operate the infrastructure and tooling that enable design, simulation, CI, and deployment for high-performance ASIC development. The role focuses on hybrid on-premise/cloud HPC clusters, observability, reproducible infrastructure control planes, and developer-facing abstractions.
This is an in-person engineering role based in San Jose focused on performance, reliability, and scale for ASIC and EDA workloads.
Experience Level
Senior — the role expects an experienced engineer; the posting specifies 8+ years of experience in infrastructure engineering, systems programming, or backend development.
Responsibilities
Deliver production-grade infrastructure and platform tooling that supports ASIC development at scale.
- Design and implement orchestration for hybrid HPC clusters to run simulation, synthesis, emulation, and massively parallel CI.
- Build a programmable infrastructure control plane for reproducibility, auditing, and rapid iteration.
- Implement workload orchestration and migration strategies across on-prem and cloud to balance performance, storage, uptime, and cost.
- Architect and operate a real-time observability stack (metrics, tracing, logs, dashboards, alerts) with streaming telemetry and LLM-based insight extraction.
- Create developer-facing tooling and abstractions to provision interactive environments (Jupyter/VS Code) and attach GPUs/high-memory nodes securely.
- Develop synthetic testing, fault-injection, and monitoring to validate infrastructure behavior under load and degraded conditions.
- Lead technical design, mentor engineers, and drive engineering best practices for infrastructure-as-software.
Requirements
Core technical skills and experience required to perform the role.
- Must-have: 8+ years in infrastructure engineering, systems programming, or backend development with production software engineering discipline.
- Must-have: Strong programming skills in Python, Go, Rust, or C++ and experience building production-grade tooling.
- Must-have: Expert knowledge of Linux, virtualization, containerization, CI/CD, and debugging/optimizing large systems.
- Must-have: Experience with infrastructure-as-code and automation tools (Terraform/OpenTofu, Ansible, Puppet, or equivalent).
- Must-have: Experience designing and operating observability stacks (Prometheus, Grafana, VictoriaMetrics, tracing, log aggregation) and using PromQL or similar.
- Must-have: Track record of solving hardware–software integration issues across bare-metal, networking, and distributed workloads.
- Nice-to-have: Experience with SLURM, Kubernetes, MaaS, Bazel, or EDA toolchain workflows (Synopsys, Cadence, Verilator).
- Nice-to-have: Hands-on experience with hybrid cloud/on-prem deployments (AWS, GCP, Azure) and workload migration strategies.
Education Requirements
Not specified.
About the Company
Company: Etched
Headquarters: San Jose, CA, United States
Etched develops purpose-built AI inference ASICs and systems optimized for transformer models, aiming to deliver significantly higher performance, lower cost, and lower latency than GPUs. The company focuses on enabling applications like real-time video generation and advanced reasoning agents, and is backed by leading investors and engineers.
