Supercomputing Engineer
EtchedJob Title
Supercomputing Engineer
Role Summary
The Supercomputing Engineer will build foundational software for cluster-scale AI compute deployments, focusing on control-plane software, system bring-up, telemetry, orchestration primitives, and performance tuning at the hardware–software boundary.
This role sits on the Supercomputing team in San Jose and requires close collaboration with hardware, firmware, kernel, and runtime engineers to enable reliable, high-performance inference infrastructure.
Experience Level
Mid-level (no explicit years-of-experience stated).
Responsibilities
Primary responsibilities include developing low-level system software, debugging at the hardware–software interface, and driving performance and reliability for rack- and cluster-scale systems.
- Architect and implement low-level control-plane software for system bring-up, configuration, and management of cluster-scale AI compute deployments
- Design and build system services that interact directly with hardware, firmware, and the OS
- Develop telemetry, logging, and tracing infrastructure to diagnose failures and guide performance improvements
- Implement orchestration primitives for managing devices, nodes, and racks
- Profile and tune performance across PCIe, memory, networking, kernel, and runtime layers
- Collaborate with hardware, firmware, kernel, and runtime teams to co-design system interfaces and behavior
Requirements
Key qualifications and technical skills required and preferred.
Must-have:
- Strong proficiency in C/C++ or Rust for low-level systems programming
- Deep understanding of Linux internals and kernel/user-space boundaries
- Experience working close to hardware: drivers, DMA, interrupts, memory management, or device control paths
- Strong debugging skills using logs, tracing, and low-level observability tools
- Effective communication and experience collaborating across hardware and software teams
Nice-to-have:
- Experience with data center orchestration technologies such as Kubernetes and Docker
- Kernel development, device drivers, or firmware-adjacent software experience
- Familiarity with PCIe, NUMA, networking, or high-speed interconnects
- Experience with tracing/profiling tools (perf, eBPF, ftrace) and failure-injection frameworks
- Background in HPC, AI infrastructure, or large-scale compute systems
Education Requirements
Not specified.
About the Company
Company: Etched
Headquarters: San Jose, CA, United States
Etched develops purpose-built AI inference ASICs and systems optimized for transformer models, aiming to deliver significantly higher performance, lower cost, and lower latency than GPUs. The company focuses on enabling applications like real-time video generation and advanced reasoning agents, and is backed by leading investors and engineers.
