Software Engineer, AI Model Enablement & Inference
FuriosaAIJob Title
Software Engineer, AI Model Enablement & Inference
Role Summary
Join the MLSys team to implement and optimize production-ready model kernels and integrations that enable efficient inference of large language models on FuriosaAI NPUs. You will develop kernel implementations in FuriosaAI's Tensor Contraction Language (TCL) and collaborate with compiler and inference engine teams to integrate models into Furiosa-LLM (FLM).
Experience Level
Mid-level. No specific years-of-experience stated.
Responsibilities
Primary responsibilities focus on analyzing models, implementing and optimizing TCL kernels, integrating them into the inference stack, and validating performance and correctness.
- Analyze model architectures and reference implementations to define implementation requirements and trade-offs; coordinate strategies with Inference Engine and Compiler teams.
- Design, implement, and optimize model-specific TCL kernels (e.g., attention, mixture-of-experts) for efficient NPU utilization.
- Integrate new models and kernels into Furiosa-LLM to enable correct and efficient inference.
- Develop reusable analysis, integration, validation, and benchmarking tooling to streamline adding new models.
- Validate kernel and model correctness on NPUs against references, investigate numerical differences, and build regression tests.
- Evaluate and adapt techniques from frameworks like vLLM and SGLang for TCL and FLM integration; document validated findings.
- Keep current with state-of-the-art model architectures and provide technical guidance on model capabilities and inference trade-offs.
Requirements
Must-have technical skills and domain knowledge are listed below; preferred skills follow.
- Must-have: Deep understanding of transformer-based LLMs, including attention variants, mixture-of-experts architectures, and KV-cache behavior.
- Must-have: Strong Python skills and hands-on experience reading, modifying, and debugging model implementations in PyTorch or comparable frameworks.
- Must-have: Hands-on experience implementing, debugging, and optimizing tensor operations or accelerator kernels with awareness of compute, memory, and numerical correctness.
- Must-have: Knowledge of LLM inference performance (prefill/decode, batching, latency–throughput trade-offs) and experience with quantitative performance evaluation.
- Must-have: Ability to read and reason about Rust or C++ code when working on inference systems.
- Must-have: Clear technical communication and cross-team collaboration skills.
- Preferred: Experience bringing up or optimizing models on GPUs, NPUs, TPUs, or other AI accelerators.
- Preferred: Familiarity with vLLM, SGLang, TensorRT-LLM, or similar inference frameworks and their scheduling/caching/model-parallel approaches.
- Preferred: Familiarity with ML compiler optimizations such as fusion, tiling, and scheduling.
- Preferred: Experience developing accelerator kernels with CUDA, Triton, or a tensor DSL; or software development in Rust, and building validation/benchmarking tools.
Education Requirements
Not specified.
About the Company
Company: FuriosaAI
Headquarters: Seoul, South Korea
FuriosaAI develops high-performance, energy-efficient AI inference hardware and software. Founded in 2017 by semiconductor and AI engineers, the company builds AI-native compute platforms to reduce AI energy and operational costs and operates globally with offices in Korea, Silicon Valley, and an R&D lab in Lisbon.
