Staff Modeling Architect
NeurophosJob Title
Staff Modeling Architect
Role Summary
Senior technical role responsible for building functional and performance models that connect production ML workloads to architecture, RTL, and software. You will own workload areas and the modeling methodology while coordinating with architects, RTL/physical design, compiler, and runtime teams.
The position focuses on hardware/software co-design for inference accelerators: converting models and papers into validated performance, power, and functional artifacts used to drive microarchitecture, ISA, memory hierarchy, and multi-chip mapping decisions.
Experience Level
Senior-level β typically 8+ years of relevant experience in hardware, functional, or performance modeling; equivalent practical experience considered.
Responsibilities
Operate across software and hardware teams to deliver models and numbers that guide design and enable software bring-up prior to silicon.
- Bring up inference workloads (transformers, MoE, attention/KV cache, SSM, hybrid models, retrieval, speech, vision, recommendation) on the model stack.
- Bind Hugging Face and PyTorch workloads to the programming model and runtime and run them on functional models.
- Co-design tiling, scheduling, ISA, SRAM/HBM hierarchy, NoC traffic, and multi-chip mapping across pipeline, tensor, and sequence parallelism.
- Run roofline and limiter analysis and design-space exploration to identify and resolve bottlenecks between compiler and hardware views.
- Develop Python energy and latency models (NumPy, Pandas, Matplotlib) covering operators, tiling, memory traffic, and optical GEMM/vector-unit time.
- Implement bit-accurate C++ functional models of optical GEMM, SRAM vector processors, dataflow engines, and HBM, including narrow arithmetic.
- Contribute to the C++ event-driven simulation kernel (coroutines, timed components, traces) and implement cycle-approximate/accurate PPA models aligned with RTL via co-simulation.
- Maintain consistency across roofline, limiter, performance models, and RTL simulation; document and arbitrate disagreements.
- Set modeling methodology for workload areas and mentor engineers in your domain.
Requirements
Must-have technical skills and experience for immediate contribution.
Must-have:- 8+ years of experience in hardware/functional/performance modeling, performance simulation, or accelerator performance analysis; graduate research may count toward this.
- Proven track record of delivering a model or study another team relied upon (architecture, compiler, customer, or silicon).
- Strong grounding in computer architecture, microarchitecture, memory systems, and AI accelerators (GPU/TPU/NPU/custom SoC).
- Modern C++ (C++17 or later) for functional models and simulation infrastructure.
- Python for models, analysis, and plotting (NumPy, Pandas, Matplotlib).
- Experience working inside or extending discrete-event, cycle-approximate, or cycle-accurate simulators (e.g., SystemC, gem5, SST, or custom kernels).
- Ability to build an LLM or accelerator workload from a model card or paper, including prefill/decode, MoE, GEMM tiling, and quantization.
- Judgment to choose appropriate methods (roofline, limiter, analytical models, trace-driven, TLM, RTL simulation) for given questions.
- Hardware/software co-design experience with compiler/runtime/ISA (MLIR, TVM, XLA, ONNX, operator fusion, graph compilers).
- Experience modifying or extending simulation kernels, or correlating models against silicon, datasheets, or measured datacenter devices.
- Familiarity with Verilator, SystemVerilog, DPI, UVM, or TLM 2.x.
- Familiarity with HBM, DRAM controllers, caches, SRAM, NoC, AXI, DMA, and scratchpad memory.
- Power modeling experience (McPAT, CACTI, or custom) and FPGA prototyping or hardware emulation.
Education Requirements
BS, MS, or PhD in Computer Engineering, Electrical Engineering, Computer Science, or equivalent practical experience. The posting explicitly allows equivalent practical experience and notes graduate research may count toward required experience; a PhD is listed as preferred.
What We Offer
Senior engineering opportunity at an early-stage startup working on photonics-based AI accelerators. Competitive compensation and benefits, 100% base health premium coverage for employees and dependents, unlimited PTO, 401(k) matching, stock options, and a suite of voluntary benefits.
About the Company
Company: Neurophos
Headquarters: Sunnyvale, California, United States
Startup developing silicon-photonic AI accelerator chips that use programmable metasurfaces and dense optical cells to perform matrix multiplications at the speed of light, targeting large-scale AI inference with much higher energy efficiency and performance than traditional electronic approaches.
