Senior Software Engineer - Kernels
d-MatrixJob Title
Senior Software Engineer - Kernels
Role Summary
Develop, optimize, and maintain software kernels and related toolchain to productize the software stack for a next-generation AI compute engine. Collaborate with compiler, ML, systems and hardware teams to map computational graphs and algorithms to the underlying hardware and drive hardware–software co-design decisions.
Hybrid role based in the Belgrade office (3–5 days per week).
Experience Level
Senior — typical background: MS +5 years industry experience or PhD +1 year. Requires significant experience with kernel development for specialized hardware and hardware–software co-design.
Responsibilities
Core responsibilities for this role include development, optimization, and cross-team collaboration focused on software kernels for AI hardware.
- Design, develop, enhance, and maintain software kernels for AI accelerators and related HW architectures.
- Map algorithms and computational graphs from ML frameworks to target hardware, ensuring correctness and performance.
- Optimize operators commonly used in ML workloads (GEMM, convolutions, BLAS, SIMD ops such as softmax, layer normalization, pooling).
- Work across the full-stack toolchain and collaborate with compiler experts to build compiler infrastructure and runtime components.
- Make engineering trade-offs for hardware–software co-design and optimize for performance, area, and power where relevant.
- Deliver scalable software deliverables within tight development timelines and participate in code and design reviews.
Requirements
Must-have technical skills and practical experience.
- Strong understanding of computer architecture, data structures, system software, and machine learning fundamentals.
- Proficient in C/C++ and Python development on Linux using standard development tools and workflows.
- Experience implementing algorithms and ML operators in high-level languages (C/C++, Python).
- Experience developing for specialized hardware (FPGAs, DSPs, GPUs, AI accelerators) and using libraries such as CUDA.
- Experience with embedded SIMD/vector processors (example: Tensilica) and SIMD programming techniques.
- Practical experience mapping computational graphs from ML frameworks to hardware and working across the full toolchain.
- Self-motivated, strong ownership, good communicator, and effective collaborator in cross-functional teams.
Nice-to-have:
- Prior startup, small-team, or incubation experience.
- Experience with ML frameworks such as TensorFlow and/or PyTorch.
- Experience with ML compilers and tools (MLIR, LLVM, TVM, Glow, etc.).
- Experience with deep learning models for CV, NLP, or recommendation systems.
- Work experience at a cloud provider or AI compute/subsystem company.
Education Requirements
MS in computer engineering, mathematics, physics, or a related technical field with 5+ years of industry experience; or PhD in computer engineering, mathematics, physics, or a related technical field with 1+ years of industry experience.
About the Company
Company: d-Matrix
Headquarters: Santa Clara, California, United States
d-Matrix is a Santa Clara–based startup developing highly programmable in-memory computing architectures and accompanying software to accelerate generative AI and other AI workloads, focusing on hardware-software co-design for cloud and edge applications.
