Senior/Staff Deep Learning Compiler Engineer
QuadricJob Title
Senior/Staff Deep Learning Compiler Engineer
Role Summary
Design and optimize the compiler stack (MLIR, TVM, LLVM) to map modern neural network frameworks to Quadric’s GPNPU for edge devices. Work closely with architecture and hardware teams to maximize performance, reduce latency, and improve memory efficiency for edge AI workloads.
Contribute to hardware-software co-design and validation on FPGA/ASIC platforms.
Experience Level
Senior - typically 8+ years of relevant experience in compiler development, deep learning systems, or software/hardware co-design.
Responsibilities
Lead implementation and optimization of compiler passes, model lowering, profiling, and hardware integration to deliver production-quality performance.
- Design, implement, and maintain compiler optimization passes targeting Quadric’s processor architecture (MLIR, Apache TVM, LLVM).
- Develop lowering pathways from PyTorch, TensorFlow, and ONNX to optimized low-level kernel code.
- Optimize models for memory throughput, latency, compute utilization, and power.
- Implement graph-level optimizations: operator fusion, layout transformations, quantization (INT8/FP16), and memory allocation strategies.
- Analyze novel model topologies (Transformers, CNNs, vision-language) and extend compiler support for new operators/primitives.
- Benchmark and profile end-to-end model performance; identify and resolve compiler bottlenecks.
- Collaborate with hardware and micro-architecture teams on ISA extensions and accelerator features.
- Develop software simulators, functional models, and test benches; participate in FPGA/ASIC bring-up and validation.
Requirements
Key technical requirements (must-have vs nice-to-have).
- Must-have: Hands-on experience developing deep learning compilers or compiler infrastructures (MLIR, Apache TVM, LLVM, XLA, TensorRT, or Glow).
- Must-have: Strong proficiency in C++ (C++14/17/20) and Python; solid fundamentals in data structures, algorithms, and object-oriented design.
- Must-have: Familiarity with modern AI/ML frameworks (PyTorch, TensorFlow, ONNX) and operator representations.
- Must-have: Solid understanding of computer systems: memory hierarchies, parallel processing, and CPU/GPU/NPU instruction execution.
- Nice-to-have: Low-level kernel optimization, SIMD/vector programming, and advanced memory allocation strategies.
- Nice-to-have: Experience with model quantization methodologies (INT8, FP8, mixed-precision) and post-training/QAT techniques.
- Nice-to-have: Prior experience on software stacks for custom AI accelerators, DSPs, or embedded architectures.
- Nice-to-have: FPGA bring-up, hardware emulation, or cycle-accurate simulator development experience.
Education Requirements
Bachelor's, Master's, or Ph.D. (BS/MS/PhD) in Computer Science, Computer Engineering, Electrical Engineering, or a related technical field.
About the Company
Company: Quadric
Headquarters: Burlingame, California, United States
Quadric is building the world’s first supercomputer designed for the real-time needs of edge devices. Founded in 2016, the company empowers developers across industries with innovative general-purpose neural processing unit (GPNPU) architecture for neural network workloads. Co-founded by technologists from MIT and Carnegie Mellon, Quadric aims to enable groundbreaking technology development.
