Performance Tools Intern
EtchedJob Title
Performance Tools Intern
Role Summary
Design and implement performance analysis and profiling tooling for a custom ML accelerator. Work with hardware, compiler, firmware, and inference engineers to collect and interpret traces, hardware counters, and memory behavior to identify bottlenecks and improve throughput and latency.
This is an in-person internship based in San Jose (Santana Row) contributing to low-level tooling for a PCIe-attached accelerator.
Experience Level
Entry-level (Internship). Suitable for students or early-career engineers; no specific years-of-experience stated.
Responsibilities
Typical tasks you will perform during the internship include:
- Build components of performance analysis and profiling infrastructure for a custom ML accelerator.
- Collect and analyze hardware counters, execution traces, and memory behavior from accelerator hardware.
- Develop low-overhead tracing of host runtime, driver activity, PCIe transfers, and accelerator execution.
- Correlate performance events across CPU, accelerator, storage, networking, and distributed workloads.
- Create analysis passes to detect memory access inefficiencies and PCIe bandwidth saturation.
- Implement visualization and timeline views that correlate CPU API calls, driver submissions, and accelerator execution units.
- Collaborate with hardware, compiler, firmware, and inference teams to iterate on tooling and benchmarks.
Requirements
Must-have skills and experience, plus a short list of nice-to-have qualifications.
- Must-have: Strong programming skills in C++ or Rust; Python experience is a plus.
- Must-have: Solid understanding of computer architecture (CPUs, GPUs/accelerators), memory hierarchies, and parallel programming.
- Must-have: Interest or experience in low-level performance analysis, profiling, and performance optimization.
- Must-have: Familiarity or interest in operating systems, compilers, firmware, drivers, or other low-level systems software.
- Must-have: Strong problem-solving skills and ability to learn quickly in a fast-paced environment.
- Nice-to-have: Direct experience developing performance analysis or debugging tools, experience with ML accelerator architectures (GPUs, TPUs), or kernel-mode driver development (Linux or Windows).
- Nice-to-have: Familiarity with performance tools such as Nsight, VTune, Perfetto, or similar profilers and tracers.
Education Requirements
Not specified.
About the Company
Company: Etched
Headquarters: San Jose, CA, United States
Etched develops purpose-built AI inference ASICs and systems optimized for transformer models, aiming to deliver significantly higher performance, lower cost, and lower latency than GPUs. The company focuses on enabling applications like real-time video generation and advanced reasoning agents, and is backed by leading investors and engineers.
