Research Engineer - CUDA Kernel Engineering
VoltaiJob Title
Research Engineer — CUDA Kernel Engineering
Role Summary
Develop and optimize CUDA kernels and integration layers to enable large-scale training, inference, and reinforcement-learning systems for AI-driven semiconductor design and verification. Work with researchers and engineers to maximize GPU efficiency across multi‑GPU deployments and deliver production-ready kernels, benchmarks, and tooling.
Expect to build performance benchmarks, CI/test suites, and contribute selected kernels and tools to open-source AI/HPC ecosystems.
Experience Level
Mid-level.
Responsibilities
Core responsibilities include:
- Design, implement, and optimize CUDA kernels for compute- and memory-bound primitives (attention, routing, graph operations, physics-inspired operators, etc.).
- Profile GPU workloads and resolve performance bottlenecks (compute, memory, interconnect, occupancy).
- Integrate custom kernels and operators into training and inference frameworks (e.g., PyTorch, Megatron, vLLM, TorchTitan).
- Develop multi‑GPU scaling strategies and handle communication (NVLink, NCCL) and memory management for large models.
- Build benchmarks, tests, and automation to validate correctness and measure performance across hardware generations.
- Collaborate with AI researchers and semiconductor domain experts to translate domain-specific algorithms into high-performance GPU code.
- Package, document, and maintain kernels and tooling for internal use and selected open-source releases.
Requirements
Key qualifications for the role. Separate must-have and nice-to-have items are listed below.
Must-have
- Practical experience writing and optimizing CUDA kernels for large-scale AI workloads.
- Strong skills in C++ and CUDA, and experience with GPU performance tools and workflows (Nsight, CUPTI, nvprof, profilers).
- Experience integrating custom CUDA ops into PyTorch or comparable training/inference frameworks.
- Familiarity with NVIDIA hardware and software stack (recent architectures, NVLink, NCCL, Triton).
- Experience with multi‑GPU scaling and inter-GPU communication patterns.
- Ability to create benchmarks, tests, and CI for kernel correctness and performance.
Nice-to-have
- Experience with Triton or other GPU code-generation frameworks.
- Background in graph reasoning, symbolic computation, or hardware simulation kernels.
- Previous contributions to open-source AI or HPC projects.
- Familiarity with reinforcement learning training systems or large-model engineering.
Education Requirements
Not specified.
About the Company
Company: Voltai
Headquarters: Palo Alto, CA, United States
Voltai develops AI-driven world models and agents to design, evaluate, and optimize physical systems—focusing on hardware, electronics, and semiconductors to enable AI-led hardware co-design, performance modeling, and cross-domain optimization.
