Principal Software Engineer, Kernels
d-MatrixJob Title
Principal Software Engineer, Kernels
Role Summary
Develop and optimize software kernels and related toolchain components for a next-generation AI compute engine. Work on mapping ML computational graphs and algorithms to hardware, collaborating closely with compiler, systems, ML software, and hardware engineering teams.
Experience Level
Senior — expects substantial industry experience. Typical guidance from the role: around 10+ years (industry) for master's-level candidates or 5+ years for PhD-level candidates.
Responsibilities
Deliver, optimize, and maintain software kernels and runtime components that enable ML workloads on specialized hardware.
- Design and implement efficient kernel implementations (GEMM, convolutions, BLAS, softmax, layer norm, pooling, etc.).
- Map computational graphs from ML frameworks to underlying hardware and optimize execution.
- Work with compiler and toolchain teams to build and integrate compiler infrastructure and runtime support.
- Optimize performance for specialized hardware (SIMD/vector processors, DSPs, GPUs, FPGAs, AI accelerators).
- Profile, benchmark, and iterate on performance, memory, and power trade-offs.
- Collaborate with hardware, mixed-signal, DSP, and CPU engineers to inform hardware–software co-design decisions.
- Deliver production-quality software on aggressive schedules and help scale deliverables across the product stack.
- Mentor peers and contribute to cross-functional design reviews and technical planning.
Requirements
Core technical skills and experience required for the role.
- Strong understanding of computer architecture, data structures, system software, and ML fundamentals.
- Proficiency in C/C++ and Python in Linux environments and familiarity with standard development tools.
- Experience implementing algorithms in high-level languages and optimizing them for performance.
- Experience developing for specialized hardware (FPGAs, DSPs, GPUs, AI accelerators) and using libraries such as CUDA.
- Experience implementing ML operators and SIMD/vectorized implementations for ML workloads.
- Experience with embedded SIMD/vector processors (example: Tensilica) or similar architectures.
- Practical experience across the full-stack toolchain (compiler, runtime, kernels) and hardware–software co-design.
- Self-motivated, strong ownership, and effective team collaboration skills.
Preferred:
- Startup or small-team experience.
- Experience with ML frameworks (TensorFlow, PyTorch) and ML compilers/tools (MLIR, LLVM, TVM, Glow).
- Experience with ML models and workloads for CV, NLP, or recommendation systems.
- Experience at a cloud provider or AI compute/subsystem company.
Education Requirements
MS in computer engineering, mathematics, physics, or a related technical field with ~10+ years of industry experience; or a PhD in computer engineering, mathematics, physics, or a related technical field with ~5+ years of industry experience.
About the Company
Company: d-Matrix
Headquarters: Santa Clara, California, United States
d-Matrix is a Santa Clara–based startup developing highly programmable in-memory computing architectures and accompanying software to accelerate generative AI and other AI workloads, focusing on hardware-software co-design for cloud and edge applications.
