Skip to main content
C

ML Systems Engineer

ChipAgents
August 27, 2026
Full-time
On-site
San Jose, California, United States
$150,000 - $350,000 USD yearly
EDA Jobs, Level - Mid-Career

Job Title

ML Systems Engineer

Role Summary

The ML Systems Engineer optimizes large language model (LLM) inference performance and efficiency for ChipAgents' agentic AI platform. The role focuses on low-level systems optimization, profiling, benchmarking, and architecting multi-node clusters for production training and inference.

You will work with research scientists and production teams to reduce latency, increase throughput, and lower inference costs for real-world semiconductor design workloads.

Experience Level

Mid-level. The role expects hands-on experience with large-scale ML systems, GPU computing, and production inference optimization; no specific years-of-experience were provided.

Responsibilities

Primary responsibilities include designing and operating high-performance inference systems and building evaluation infrastructure.

  • Design, deploy, and optimize multi-node LLM inference clusters to maximize throughput and minimize latency for production workloads.
  • Implement, benchmark, and validate concrete inference optimizations and batching strategies.
  • Profile and analyze systems-level bottlenecks across GPU kernel execution, memory bandwidth, and inter-node communication.
  • Build robust evaluation harnesses and benchmarking frameworks measuring accuracy, throughput, latency, and resource use across parallelism strategies.
  • Integrate new model architectures and optimizations from research into production inference infrastructure.
  • Investigate and apply emerging techniques from research and open-source projects to improve inference performance.

Requirements

Must-have technical skills and competencies; a short list of preferred additions follows.

  • Experience with large-scale ML systems, GPU computing, or high-performance inference optimization.
  • Strong proficiency in Python and C++; hands-on experience with CUDA for performance work.
  • Practical experience with inference frameworks (vLLM, SGLang, PyTorch, or similar).
  • Deep understanding of GPU architecture, memory hierarchies, and parallel computing paradigms.
  • Production experience deploying and optimizing LLMs: model serving, batching strategies, distributed inference, and quantization.
  • Strong systems-level debugging and profiling skills across stack layers from CUDA kernels to application logic.
  • Nice-to-have: familiarity with distributed computing frameworks (Ray), multi-node training/inference orchestration, and additional profiling tools.
  • Self-directed problem solver comfortable with ambitious optimization challenges.

Education Requirements

B.S., M.S., or PhD in Computer Science, Electrical Engineering, or a related field β€” or equivalent practical experience.


About the Company

Company: ChipAgents

Headquarters: San Jose, CA, United States

ChipAgents is a Series A startup developing agentic AI workflows for chip design and verification, leveraging generative AI and large language models to assist RTL design, simulation, and EDA workflows for semiconductor and cloud customers.

ChipAgents logo

Date Posted: 2026-08-27