Skip to main content
Etched logo

Applied AI Engineer, Kernel Performance

Etched
August 27, 2026
Full-time
On-site
San Jose, California, United States
$150,000 - $225,000 USD yearly
EDA Jobs, Level - Mid-Career

Job Title

Applied AI Engineer, Kernel Performance

Role Summary

Build AI systems that autonomously convert newly released model architectures into correct, production-ready kernels and model mappings optimized for Etched hardware. Work across tooling, agents, profiling, and kernel engineering to accelerate discovery of high-performance implementations.

Collaborate with architecture and runtime teams, run high-throughput experiments, and ship model-generated improvements that measurably improve end-to-end system performance.

Experience Level

Mid-level. No specific years-of-experience requirement was stated.

Responsibilities

Own and improve the system that generates, verifies, profiles, and ships performant kernels and mappings for new models on Etched hardware.

  • Own the end-to-end system converting model architectures into verified, production-ready kernels and mappings.
  • Build agentic systems that design experiments, generate implementations, profile results, diagnose bottlenecks, and iterate toward better proposals.
  • Design and run evaluations for correctness, numerical stability, latency, and efficiency.
  • Turn profiler traces, simulations, hardware counters, and expert judgment into structured learning signals.
  • Curate proprietary datasets from experiment trajectories, demonstrations, counterexamples, and production outcomes.
  • Build reproducible, high-throughput experiment infrastructure and observability for interpretable results.
  • Ship model-generated improvements to production and quantify their impact on system-level performance.
  • Partner with architecture teams to shape abstractions and roadmap decisions.
  • Continuously evaluate new model releases and deploy the best-performing models into the optimization loop.

Requirements

Must-have:

  • Proven track record solving hard engineering problems across stacks and domains; able to learn quickly in unfamiliar areas.
  • Proficiency with Python and familiarity reading and modifying low-level code (C/C++, CUDA, or similar); able to debug and direct AI to produce correct code.
  • Kernel experience: have written or tuned kernels and explain the mechanisms and performance impact of optimizations you've shipped.
  • Experience using AI tools and agents as part of development and experimentation workflows.
  • Experience with performance analysis using profilers, traces, simulators, or hardware counters.
  • Experience shipping production code and measuring performance impact.
  • Must be able to work on-site in San Jose (Santana Row) or relocate; relocation support is available.

Nice-to-have:

  • First-principles experience with accelerator performance: memory hierarchy, data movement, parallelism, synchronization, and low-precision computation.
  • Hands-on experience building and shipping LLM-based agents or AI tooling in production (context engineering, tool integration, orchestration, failure analysis).
  • Eval-driven mindset, fine-tuning or post-training experience, RAG over proprietary data, or multi-agent orchestration experience.
  • Experience with hardware-software co-design or systems-level kernel optimization platforms.

Education Requirements

Not specified.


About the Company

Company: Etched

Headquarters: San Jose, CA, United States

Etched develops purpose-built AI inference ASICs and systems optimized for transformer models, aiming to deliver significantly higher performance, lower cost, and lower latency than GPUs. The company focuses on enabling applications like real-time video generation and advanced reasoning agents, and is backed by leading investors and engineers.

Etched logo

Date Posted: 2026-08-27