ML Systems Engineer — Inference Acceleration
AragoJob Title
ML Systems Engineer — Inference Acceleration
Role Summary
Optimize execution and serving of modern AI models on Arago's custom optical+CMOS accelerator. Work across kernels, model execution, multi-device distribution, runtime, and inference serving, and collaborate with hardware, compiler, and runtime teams to shape the software stack.
Experience Level
Mid-level. Years of experience not specified.
Responsibilities
Key responsibilities include:
- Analyze AI workloads and identify kernel-, runtime-, memory-, and system-level bottlenecks on Arago's accelerator.
- Develop and optimize custom kernels, fused operators, and execution strategies to maximize device utilization.
- Design efficient mappings of models and operators across multiple devices, including communication and synchronization strategies.
- Implement inference-serving techniques such as continuous batching, paged KV caches, prefix/context caching, chunked prefill, and prefill/decode interleaving.
- Build profiling, benchmarking, and performance-analysis infrastructure for kernels, full models, and serving workloads.
- Collaborate with hardware, compiler, and runtime teams to co-design software abstractions and influence future hardware features.
Requirements
Must-have skills and experience:
- Strong experience in high-performance ML inference, GPU/accelerator programming, or ML systems engineering.
- Deep understanding of computer architecture, accelerator/GPU execution models, memory hierarchies, parallelism, and performance bottlenecks.
- Experience developing and optimizing custom kernels using CUDA, Triton, ROCm/HIP, or equivalent low-level programming environments.
- Experience with operator fusion, tiling, scheduling, data movement optimization, graph execution, and profiling of compute- and memory-bound workloads.
- Knowledge of distributed model execution (tensor, pipeline, sequence, and/or expert parallelism) and communication/computation overlap.
- Hands-on experience with inference-serving systems (vLLM, SGLang, TensorRT-LLM, or equivalent), including KV-cache management and batching strategies.
- Proficient in C++ and Python; comfortable working on a custom accelerator stack where compiler, runtime, kernels, and abstractions are actively developed.
- English proficiency at a working level.
Nice-to-have:
- Exposure to or experience with MLIR and MLIR dialects.
- Prior work on optical or proprietary accelerator platforms.
Education Requirements
Not specified.
About the Company
Company: Arago
Arago is an AI and computer hardware startup (founded 2024) developing AI-driven semiconductor and photonics solutions. The company brings together engineers and researchers in photonics, electronics, software, mathematics, and machine learning to accelerate prototype-to-silicon development across hubs in France, North America, and Israel.
