Algorithm - AI System Engineer
FuriosaAIJob Title
Algorithm - AI System Engineer
Role Summary
Implement and validate next-generation serving-system concepts on Furiosa's NPU hardware, focusing on proof-of-concept (POC) and prototyping that can feed into productization with software teams.
The role emphasizes low-level implementation, performance verification on real devices, and turning implementation lessons into new research and optimization proposals.
Experience Level
Mid-level β role expects practical experience in AI inference systems and accelerator programming. (Source indicates 3+ years of related experience or equivalent.)
Responsibilities
Work across research, implementation, and verification to prove and refine serving-system ideas on NPU hardware.
- Implement core elements of next-generation serving systems (e.g., Attention-FFN Disaggregation, KV cache reuse) as NPU-based POCs and prototypes.
- Explore and implement compression techniques such as quantization and KV cache compression in an NPU environment and validate their impact.
- Use the low-level programming stack and kernel interfaces to extract hardware performance and identify bottlenecks.
- Design, run, and analyze experiments and simulations; validate ideas on real hardware.
- Convert implementation findings into new research questions, optimization proposals, and actionable improvements for product teams.
- Communicate technical trade-offs and lead cross-team collaboration with software and product stakeholders.
Requirements
Must-have technical skills and experience required for the role. Preferred items are listed separately.
- Understanding of LLM inference mechanics (attention, KV cache, prefill/decode, batching).
- Knowledge or hands-on experience with AI inference systems such as vLLM, SGLang, or TensorRT-LLM.
- Experience with accelerator programming (CUDA, Triton, or similar).
- Proven ability to solve ill-defined problems through experiments and systematic debugging.
- Clear technical communication and ability to lead collaboration across teams.
Nice-to-have / preferred:
- Understanding of low-level system concepts: multi-threading, memory management, performance optimization.
- Familiarity with multi-chip parallelism (tensor/pipeline parallelism) and distributed inference.
- Experience with model quantization, KV cache optimization, or other compression techniques.
- Prior work in deep-learning acceleration or inference-platform teams.
- Familiarity with AI coding tools for development and research workflows.
Education Requirements
Minimum of 3 years of relevant practical experience or equivalent practical experience; no specific degree or field of study is required or specified.
About the Company
Company: FuriosaAI
Headquarters: Seoul, South Korea
FuriosaAI develops high-performance, energy-efficient AI inference hardware and software. Founded in 2017 by semiconductor and AI engineers, the company builds AI-native compute platforms to reduce AI energy and operational costs and operates globally with offices in Korea, Silicon Valley, and an R&D lab in Lisbon.
