Skip to main content
Etched logo

Inference Intern

Etched
August 27, 2026
Internship
On-site
San Jose, California, United States
SoC Architecture Jobs, Level - Entry or Early Career

Job Title

Inference Intern

Role Summary

Paid 12-week in-person internship based in San Jose, CA working on inference-focused AI accelerator architecture and software. Interns will develop and optimize compute architectures, performance models, and runtime components to improve inference throughput and latency.

You'll work with engineering teams to port models, build runtime and tooling, profile performance, and influence co-design between hardware instructions and model operations.

Experience Level

Entry-level internship. Open to candidates for Fall '26, Spring '27, and Summer '27 internship terms.

Responsibilities

Primary contributions expected during the internship:

  • Port state-of-the-art models to the company architecture; build programming abstractions and tests to accelerate iteration.
  • Assist in building, enhancing, and scaling the runtime: multi-node inference, intra-node execution, state management, and robust error handling.
  • Optimize routing and communication layers, including collective operations.
  • Use performance profiling and debugging tools to identify bottlenecks and correctness issues.
  • Co-design HW instructions and model operations to maximize model performance.
  • Implement high-performance software components for the Model Toolkit.

Requirements

Must-have technical skills and experience:

  • Proficiency in Python and C++.
  • Understanding of performance-sensitive or complex distributed software systems (examples: Linux internals, accelerator architectures such as GPUs/TPUs, compilers, or high-speed interconnects like NVLink/InfiniBand).
  • Experience porting applications to non-standard accelerator hardware or custom hardware platforms.
  • Familiarity with transformer model architectures and inference serving stacks (e.g., vLLM).
  • Experience with performance profiling and debugging tools for identifying bottlenecks.

Nice-to-have

  • Proficiency in Rust.
  • Experience with low-latency, high-performance networking (kernel-level and user-space).
  • Deep knowledge of distributed systems concepts (consensus, consistency, communication patterns).
  • Experience with Mixture-of-Experts (MoE) models, SIMD optimizations, or accelerator-specific optimizations.
  • Familiarity with PyTorch or JAX.
  • Participation in math competitions (AIME, AMC) noted as a plus.

Education Requirements

Progress toward a Bachelor’s, Master’s, or PhD degree in Computer Science, Computer Engineering, Applied Mathematics, or a related field.


About the Company

Company: Etched

Headquarters: San Jose, CA, United States

Etched develops purpose-built AI inference ASICs and systems optimized for transformer models, aiming to deliver significantly higher performance, lower cost, and lower latency than GPUs. The company focuses on enabling applications like real-time video generation and advanced reasoning agents, and is backed by leading investors and engineers.

Etched logo

Date Posted: 2026-08-27