Staff Software Engineer, SIMD Kernels
d-MatrixJob Title
Staff Software Engineer, SIMD Kernels
Role Summary
Join the SIMD Kernels team to productize the software stack for a next-generation AI compute engine. The role focuses on implementing, optimizing, and maintaining high-performance ML operator kernels and developer-facing SDK components for specialized hardware.
Primary location is Santa Clara, CA (headquarters) or regional offices; remote candidates will be considered.
Experience Level
Senior — requires 5+ years of industry experience (per minimum qualifications).
Responsibilities
You will develop and ship performant ML operator kernels and SDK features, work across the hardware–software stack, and analyze and improve runtime performance.
- Design, implement, and optimize software kernels for ML operators (e.g., softmax, layer normalization, activation functions).
- Map algorithms and framework computational graphs onto target hardware architectures.
- Develop tooling and SDK components that make the stack intuitive for developers and enable performance analysis.
- Work across the full toolchain, including compiler and runtime interactions, to meet performance and correctness goals.
- Collaborate with hardware and software teams to resolve co‑design tradeoffs and scale solutions.
- Deliver high-quality, maintainable code in a fast-paced development environment.
Requirements
Must-have technical skills and experience for immediate contribution.
- Strong understanding of computer architecture, data structures, system software, and machine learning fundamentals.
- Proficient in C/C++ and Python development on Linux and familiar with standard development tools.
- Experience implementing algorithms in C/C++ and Python for specialized hardware (FPGAs, DSPs, GPUs, AI accelerators) and using libraries such as CUDA.
- Experience implementing ML operators and primitives: GEMMs, convolutions, softmax, layer normalization, pooling, etc.
- Proven ability to work independently, take ownership, and lead technical efforts within a team.
Preferred:
- Startup or small-team experience.
- Experience with ML frameworks (TensorFlow, PyTorch) and ML compilers/tools (MLIR, LLVM, TVM, Glow).
- Experience with deep learning models for CV, NLP, or recommendation workloads.
- Development experience for embedded SIMD/vector processors (e.g., Tensilica) or work at cloud/AI compute companies.
Education Requirements
MS or PhD in computer engineering, mathematics, physics, or a related field; role specifies 5+ years of industry experience. (Degree requirement moved here from other sections.)
About the Company
Company: d-Matrix
Headquarters: Santa Clara, California, United States
d-Matrix is a Santa Clara–based startup developing highly programmable in-memory computing architectures and accompanying software to accelerate generative AI and other AI workloads, focusing on hardware-software co-design for cloud and edge applications.
