Senior Software Engineer, Runtime
FuriosaAIJob Title
Senior Software Engineer, Runtime
Role Summary
Design and implement the low-level runtime stack that drives FuriosaAI's NPU hardware, spanning device driver interfaces, DMA-based I/O, kernel execution scheduling, multi-node inference, and embedded firmware.
This role focuses on maximizing inference throughput and minimizing latency across the full software/hardware stack for production AI inference systems.
Experience Level
Senior β typically requires 3+ years of relevant industry experience in systems programming (Rust, C, or C++).
Responsibilities
The role owns implementation and performance tuning of the runtime and firmware that control NPU hardware.
- Develop low-level runtime for DMA-based I/O operations and kernel execution scheduling to maximize throughput and minimize end-to-end latency.
- Build and optimize asynchronous execution pipelines to orchestrate data movement and compute across NPU hardware.
- Implement communication primitives for multi-node inference, including RDMA-style low-latency, high-bandwidth transfers.
- Develop embedded firmware (PERT) for the NPU's integrated ARM core to manage on-device scheduling, synchronization, and hardware resource control.
- Profile and tune system-level performance from firmware to user-space to eliminate bottlenecks in real-world inference workloads.
Requirements
Key must-have skills and experience.
- 3+ years of relevant industry experience in systems programming using Rust, C, or C++.
- Solid understanding of computer architecture fundamentals: memory hierarchy, cache coherency, OS concepts, DMA, interrupts, and MMIO.
- Strong communication skills; able to gather requirements and drive technical alignment across teams.
Nice-to-have:
- Experience in low-latency runtime systems, embedded firmware development, or high-performance I/O for accelerator hardware.
- Experience with DMA engines, scatter-gather I/O, zero-copy transfers, RDMA, or high-performance networking.
- Experience with embedded ARM firmware (bare-metal or lightweight RTOS) and CUDA low-level runtime internals (CUDA Graphs, streams).
- Kernel-level performance optimization experience (Linux kernel modules, eBPF, perf, ftrace) and profiling on accelerator/SoC platforms.
- Understanding of deep learning inference workloads and hardware execution characteristics.
Education Requirements
BS in Computer Science, Engineering, or a related field, or equivalent practical experience (explicitly stated). No other formal degree or certification requirements specified.
About the Company
Company: FuriosaAI
Headquarters: Seoul, South Korea
FuriosaAI develops high-performance, energy-efficient AI inference hardware and software. Founded in 2017 by semiconductor and AI engineers, the company builds AI-native compute platforms to reduce AI energy and operational costs and operates globally with offices in Korea, Silicon Valley, and an R&D lab in Lisbon.
