Runtime Engineer
MatXJob Title
Runtime Engineer
Role Summary
Build and own the host-side runtime stack for MatX's custom silicon used for LLM training and inference. The role sits at the HW/SW co-design boundary and implements the device-host contract, Python interop, executable formats, and serving infrastructure that compilers and ML workloads rely on.
Experience Level
Mid-level (title indicates a mid-career role; no explicit years-of-experience stated).
Responsibilities
Primary responsibilities include implementing and operating the host runtime and the interfaces that connect compilers, kernels, and ML workloads to the device.
- Develop the host-side interface library: device memory management, DMA, streams/events, synchronization primitives.
- Define and evolve the executable format and compiler→runtime contract, including versioning and quantization/weight layouts.
- Design the custom-kernel ABI and implement host-side marshaling (DLPack, buffer protocol, NumPy) to move tensors to/from the device.
- Build Python bindings (PyO3) and maintain a C-ABI shim for alternative integrations.
- Implement LLM inference serving features: paged KV cache, continuous batching, request scheduling, token streaming, and cluster orchestration primitives.
- Bring up interconnect topology from host, implement failure detection and clean teardown for stop-restructure-resume recovery across racks.
- Design and expose profiler and debugger interfaces (perf counters, traces, Python-facing surfaces) and meet targets for runtime overhead and serving throughput.
Requirements
Must-have technical skills and constraints; bonus skills listed under "Nice-to-have."
- Strong systems programming experience in Rust, C, C++, or Go, including memory management, allocator concepts, and FFI/ABI work.
- Production experience building Python interop layers (PyO3, ctypes, pybind11, or equivalent C-ABI bridges).
- Experience designing and maintaining API/ABI contracts across teams (versioning and change discipline).
- Hands-on experience with at least one accelerator programming model (CUDA, ROCm, oneAPI Level Zero, TPU, or similar) sufficient to reason about device memory, asynchronous execution, and kernel launches.
- Familiarity with ML systems (training and inference loops, tensor layouts, collectives); research depth is not required.
- Ability to work on-site in Mountain View Tuesdays–Thursdays and authorization to work in the United States. Role requires compliance with U.S. export-control restrictions.
Nice-to-have
- LLM inference internals (vLLM, TensorRT-LLM, SGLang) — paged attention and scheduler design.
- Deep Rust expertise (proc macros, unsafe/soundness reasoning, advanced lifetime/trait work).
- Custom allocator design (slab, paged, arena) or other low-level memory systems work.
- ML framework integration (PyTorch custom backends, JAX/XLA, ONNX Runtime).
- Profiler/tracing infrastructure experience (Perfetto, Nsight, or custom stacks).
- Driver-adjacent or kernel-bypass experience, and prior new-silicon bring-up.
Education Requirements
Not specified.
About the Company
Company: MatX
Headquarters: Mountain View, California, USA
MatX specializes in creating faster chips for large language models (LLMs), focusing on innovative hardware and software solutions. The company fosters a collaborative and supportive work environment, welcoming candidates of all experience levels. Their approach prioritizes deep understanding and consideration of novel methods to drive efficiency and performance in their projects, particularly in silicon design and related engineering roles.
