Software Engineer, Acceleration Kernel Development
TenstorrentJob Title
Software Engineer, Acceleration Kernel Development
Role Summary
Develop low-level, performance-critical software that implements and optimizes compute kernels for machine learning and high-performance workloads. Work on the intersection of software and hardware performance with a team of ML and hardware engineers.
This is a hybrid role based out of Toronto, ON.
Experience Level
Mid-level. Candidates at various experience levels will be considered during the interview process; specific level and offer will be based on assessed skills and experience.
Responsibilities
The primary responsibilities focus on designing, implementing, optimizing, and maintaining low-level kernels and the software stack that runs them.
- Design and implement high-performance compute kernels in C/C++ for parallel ML workloads.
- Analyze and tune instruction-level performance across latency, memory, and bandwidth dimensions.
- Profile, debug, and optimize code using tooling to reach strict performance targets.
- Integrate kernel optimizations into ML frameworks and training pipelines.
- Collaborate closely with ML engineers and hardware engineers to validate and deploy optimizations.
- Maintain, test, and document a reliable low-level software stack suitable for production workloads.
Requirements
Must-have technical skills and constraints for the role.
- Strong software engineering skills in C and C++; ability to write efficient, low-level code.
- Proven experience building and optimizing compute kernels for parallel ML or other high-performance workloads.
- Experience with performance analysis, profiling, and debugging at instruction and microarchitectural levels.
- Collaborative mindset; experience working cross-functionally with ML and hardware teams.
- Ownership of reliability, testing, and maintenance for performance-sensitive code.
- Employment contingent on eligibility to access U.S. export-controlled technology; offers may depend on citizenship, permanent residency, or ability to obtain an export license.
Nice-to-have:
- Experience with RISC-V or other modern ISAs, assembly, SIMD/vector intrinsics, or compiler backends.
- Familiarity integrating kernels into ML frameworks (for example, PyTorch or similar).
- Background in parallel algorithms, memory hierarchy optimization, or domain-specific accelerators.
Education Requirements
Not specified.
About the Company
Company: Tenstorrent
Headquarters: Austin, Texas, United States
Tenstorrent is a technology company focused on designing innovative computing solutions. They are known for their expertise in the development of advanced hardware, including ASICs and SoCs, aimed at enhancing performance and efficiency in various applications.
