Skip to main content
FuriosaAI logo

Senior Software Engineer, Inference Engine (Platform Software)

FuriosaAI
August 27, 2026
Full-time
On-site
Seoul, KR
Other Semiconductor Jobs, Level - Senior

Job Title

Senior Software Engineer, Inference Engine (Platform Software)

Role Summary

Develop and optimize a high-performance inference engine for large and multimodal language models running on FuriosaAI NPUs. Work closely with compiler and hardware teams to co-design execution and maximize throughput, latency, and memory efficiency. Research and integrate state-of-the-art inference optimization techniques into production serving software.

Experience Level

Senior — the posting expects experienced engineers. The role's minimum experience requirement is at least 3 years of relevant industry experience or equivalent practical experience.

Responsibilities

Primary responsibilities include designing, implementing, and optimizing the production inference engine and related distributed serving capabilities.

  • Design and implement next-generation inference engine for large and multimodal language models, optimized for throughput, latency, and memory efficiency.
  • Implement advanced inference optimizations such as speculative decoding, KV-cache management, tensor/model parallelism, memory-efficient execution, and scheduling.
  • Design and develop distributed and scalable inference features including prefill–decode (PD), encode–prefill–decode (EPD), disaggregated speculative decoding, and hierarchical/external KV-cache storage.
  • Collaborate with Compiler and Hardware teams to co-design execution for FuriosaAI NPUs to improve system-level throughput, latency, and memory utilization.
  • Research, evaluate, and integrate state-of-the-art inference optimizations and features from LLM serving frameworks into production systems.

Requirements

Summary of must-have skills and desirable additions.

Must-have

  • Proficiency in Rust or C++.
  • Knowledge of deep learning, large language models, and generative AI models.
  • Strong problem-solving and data-analysis skills.
  • Effective communication and collaboration skills for cross-team work.

Nice-to-have

  • Experience building inference serving systems for large models (batching, scheduling, caching, load balancing).
  • Deep understanding of performance optimization in systems.
  • Experience with C++/CUDA or Triton kernel development.
  • Contributions to open-source inference frameworks (for example vLLM, SGLang, TensorRT-LLM).

Education Requirements

BS degree in Computer Science, Engineering, or a related field, or equivalent practical experience. The posting explicitly allows equivalent practical experience in lieu of a degree.


About the Company

Company: FuriosaAI

Headquarters: Seoul, South Korea

FuriosaAI develops high-performance, energy-efficient AI inference hardware and software. Founded in 2017 by semiconductor and AI engineers, the company builds AI-native compute platforms to reduce AI energy and operational costs and operates globally with offices in Korea, Silicon Valley, and an R&D lab in Lisbon.

FuriosaAI logo

Date Posted: 2026-08-26