Senior Software Engineer, Inference Engine (Platform Software)
FuriosaAIJob Title
Senior Software Engineer, Inference Engine (Platform Software)
Role Summary
Develop and optimize a high-performance inference engine for large and multimodal language models running on FuriosaAI NPUs. Work closely with compiler and hardware teams to co-design execution and maximize throughput, latency, and memory efficiency. Research and integrate state-of-the-art inference optimization techniques into production serving software.
Experience Level
Senior — the posting expects experienced engineers. The role's minimum experience requirement is at least 3 years of relevant industry experience or equivalent practical experience.
Responsibilities
Primary responsibilities include designing, implementing, and optimizing the production inference engine and related distributed serving capabilities.
- Design and implement next-generation inference engine for large and multimodal language models, optimized for throughput, latency, and memory efficiency.
- Implement advanced inference optimizations such as speculative decoding, KV-cache management, tensor/model parallelism, memory-efficient execution, and scheduling.
- Design and develop distributed and scalable inference features including prefill–decode (PD), encode–prefill–decode (EPD), disaggregated speculative decoding, and hierarchical/external KV-cache storage.
- Collaborate with Compiler and Hardware teams to co-design execution for FuriosaAI NPUs to improve system-level throughput, latency, and memory utilization.
- Research, evaluate, and integrate state-of-the-art inference optimizations and features from LLM serving frameworks into production systems.
Requirements
Summary of must-have skills and desirable additions.
Must-have
- Proficiency in Rust or C++.
- Knowledge of deep learning, large language models, and generative AI models.
- Strong problem-solving and data-analysis skills.
- Effective communication and collaboration skills for cross-team work.
Nice-to-have
- Experience building inference serving systems for large models (batching, scheduling, caching, load balancing).
- Deep understanding of performance optimization in systems.
- Experience with C++/CUDA or Triton kernel development.
- Contributions to open-source inference frameworks (for example vLLM, SGLang, TensorRT-LLM).
Education Requirements
BS degree in Computer Science, Engineering, or a related field, or equivalent practical experience. The posting explicitly allows equivalent practical experience in lieu of a degree.
About the Company
Company: FuriosaAI
Headquarters: Seoul, South Korea
FuriosaAI develops high-performance, energy-efficient AI inference hardware and software. Founded in 2017 by semiconductor and AI engineers, the company builds AI-native compute platforms to reduce AI energy and operational costs and operates globally with offices in Korea, Silicon Valley, and an R&D lab in Lisbon.
