Staff Deep Learning Compiler Engineer
QuadricJob Title
Staff Deep Learning Compiler Engineer
Role Summary
Design and optimize the deep learning compiler stack (MLIR, TVM, LLVM) to map modern neural network frameworks to Quadric's General-Purpose Neural Processing Unit (GPNPU). The role focuses on compiler optimization, graph transformations, code generation, and hardware–software co-design to maximize performance, latency, and memory efficiency for edge AI workloads.
Experience Level
Senior - typically requires 8+ years of experience in deep learning compilers, compiler infrastructure, or closely related systems and compiler engineering.
Responsibilities
Deliver compiler features and optimizations that enable high-performance execution of neural networks on Quadric hardware, and collaborate with architecture and validation teams.
- Design, implement, and maintain compiler optimization passes targeting Quadric's processor using frameworks such as MLIR, Apache TVM, or LLVM.
- Develop lowering paths from high-level frameworks (PyTorch, TensorFlow, ONNX) to optimized low-level kernel code and code generation backends.
- Implement graph-level optimizations: operator fusion, layout transforms, quantization support (INT8/FP16), and memory allocation strategies.
- Benchmark and profile end-to-end model performance; identify and resolve compiler bottlenecks to improve throughput and latency.
- Collaborate with hardware and microarchitecture teams to define ISA extensions, hardware acceleration features, and compiler requirements.
- Develop software simulators, functional models, and test benches; participate in FPGA/ASIC bring-up and validation to verify correctness and performance.
Requirements
Concise list of required and preferred technical qualifications.
- Must-have: 8+ years hands-on experience developing deep learning compilers or compiler infrastructures (examples: MLIR, Apache TVM, LLVM, XLA, TensorRT, Glow).
- Must-have: Strong proficiency in C++ (C++14/17/20) and Python; solid fundamentals in data structures, algorithms, and object-oriented design.
- Must-have: Familiarity with modern AI/ML frameworks and operator representations (PyTorch, TensorFlow, ONNX).
- Must-have: Deep understanding of computer systems, memory hierarchies, parallel processing, and instruction execution on CPU/GPU/NPU.
- Nice-to-have: Experience with low-level kernel optimization, SIMD/vector programming, and advanced memory allocation strategies.
- Nice-to-have: Knowledge of model quantization (INT8, FP8, mixed-precision), post-training/QAT techniques, or experience with custom AI accelerators/DSPs.
- Nice-to-have: Prior work on FPGA bring-up, hardware emulation, cycle-accurate simulators, or embedded architecture software stacks.
Education Requirements
BS, MS, or Ph.D. in Computer Science, Computer Engineering, Electrical Engineering, or a related technical field.
About the Company
Company: Quadric
Headquarters: Burlingame, California, United States
Quadric is building the world’s first supercomputer designed for the real-time needs of edge devices. Founded in 2016, the company empowers developers across industries with innovative general-purpose neural processing unit (GPNPU) architecture for neural network workloads. Co-founded by technologists from MIT and Carnegie Mellon, Quadric aims to enable groundbreaking technology development.
