Senior Staff Software Engineer - Kernels
d-MatrixJob Title
Senior Staff Software Engineer - Kernels
Role Summary
Develop, optimize, and maintain software kernels and runtime components for a next-generation AI compute engine. Work within the software team to map algorithms and computational graphs produced by ML frameworks to target hardware, and drive performance and scalability across the hardware–software stack.
This is a hybrid role based in Belgrade, Serbia (onsite 3–5 days/week) collaborating closely with compiler, ML, systems, mixed-signal, DSP, and CPU engineering teams.
Experience Level
Senior-level. The role expects a seasoned engineer with substantial industry experience and demonstrated leadership in kernel, compiler, or hardware-software co-design work.
Responsibilities
Primary responsibilities include:
- Design, implement, and optimize software kernels for AI workloads on specialized hardware.
- Map computational graphs from ML frameworks to target architectures and verify correctness and performance.
- Collaborate with compiler engineers to build and extend compiler and runtime infrastructure.
- Profile and tune operators (e.g., GEMM, convolutions, BLAS, softmax, layer norm, pooling) for target hardware.
- Develop and maintain software deliverables under tight schedules and scale implementations for production.
- Work across the full toolchain and trade off hardware/software design decisions to meet performance, area, and power goals.
- Partner with ML, systems, and hardware teams to validate end-to-end functionality and performance.
Requirements
Must-have technical skills and experience:
- Strong understanding of computer architecture, data structures, system software, and machine-learning fundamentals.
- Proficient in C/C++ and Python development on Linux using standard development tools.
- Experience implementing algorithms in C/C++ and Python and optimizing them for target hardware.
- Experience targeting specialized hardware (FPGAs, DSPs, GPUs, AI accelerators) and using relevant libraries (e.g., CUDA).
- Experience implementing ML operators and SIMD-based implementations for performance-critical kernels.
- Experience with embedded SIMD/vector processors (example: Tensilica) or similar architectures.
- Self-motivated team player with strong ownership and leadership skills.
Nice-to-have:
- Prior startup, small-team, or incubation experience.
- Experience with ML frameworks such as TensorFlow or PyTorch.
- Experience with ML compilers and related toolchains (MLIR, LLVM, TVM, Glow, etc.).
- Experience with ML models for CV, NLP, or recommendation and production deployment at scale.
- Experience working at cloud providers or AI compute/subsystem companies.
Education Requirements
MS in computer engineering, mathematics, physics, or a related technical field with 10+ years of industry experience; OR PhD in computer engineering, mathematics, physics, or a related technical field with 1+ years of industry experience. Fields referenced: computer engineering, math, physics, or related degrees.
About the Company
Company: d-Matrix
Headquarters: Santa Clara, California, United States
d-Matrix is a Santa Clara–based startup developing highly programmable in-memory computing architectures and accompanying software to accelerate generative AI and other AI workloads, focusing on hardware-software co-design for cloud and edge applications.
