Skip to main content
Neurophos logo

Senior Modeling Architect, Performance Benchmarking

Neurophos
September 04, 2026
Full-time
On-site
Austin, Texas, United States
$210,000 - $250,000 USD yearly
EDA Jobs, Level - Senior

Job Title

Senior Modeling Architect, Performance Benchmarking

Role Summary

Owner of performance and energy benchmarking for the T100 optical inference accelerator. Produce repeatable, auditable measurements across modeling fidelities (analytic models, RTL simulation) and measured runs on competitor GPUs and accelerators.

Work on the Architecture and Modeling team to define and maintain measurement methodology, benchmark harnesses, workloads, and reporting used to drive architecture and product decisions.

Experience Level

Senior-level. Typically requires 5+ years of experience in GPU performance engineering, accelerator benchmarking, HPC performance measurement, or ML systems measurement.

Responsibilities

Deliver consistent, reproducible performance and energy numbers and maintain the tooling and artifacts that justify them.

  • Own performance and energy metrics used by architecture, product, and leadership; ensure consistency across modeling fidelities and measured hardware.
  • Produce results for identical workloads across roofline/limiter analysis, architecture models, RTL simulation, and measured competitor hardware.
  • Define and lock workload parameters (model/application, sequence length, batch, precision, prefill vs decode, parallelism) across fidelities.
  • Bring up inference workloads from Hugging Face, PyTorch, papers, and vendor stacks (vLLM, SGLang, TensorRT-LLM, Triton), including dense and MoE transformers and non-LLM workloads that map to the accelerator.
  • Measure competing GPUs/accelerators end-to-end: manage cloud or lab accounts, images, drivers, and run recipes.
  • Report TTFT, inter-token latency, tokens/sec, tokens/sec/watt, and energy using available instrumentation (nvidia-smi, DCGM, power capping, or equivalent).
  • Document discrepancies between RTL simulation, performance models, and measured results; attach configs, logs, and assumptions for reproducibility.
  • Maintain a reviewed internal benchmark suite and keep internal-only results separate from any externally cleared claims.

Requirements

Must-have technical skills and hands-on experience required to perform the role.

  • Proven track record building and operating benchmark harnesses that produce measured results on real GPUs or accelerators, including turning model cards/papers into runnable benchmarks.
  • Hands-on experience with roofline analysis, limiter analysis, or analytical performance modeling.
  • GPU performance analysis experience with profilers (NVIDIA Nsight Systems/Nsight Compute or equivalent) covering HBM-bound vs compute-bound analysis, precisions (FP16, BF16, FP8, INT8), and batching strategies.
  • Working knowledge of LLM inference stacks (Hugging Face, vLLM, SGLang, TensorRT-LLM) including prefill vs decode, continuous batching, and MoE.
  • Proficiency in Python for harnesses, parsing, and plotting; comfortable working in Linux.
  • Cloud GPU operations experience on AWS, GCP, or Azure: containers, instance types, drivers, quotas, and cost management.
  • Experience operating lab or cloud accounts and owning end-to-end measurement runs.

Nice-to-have:

  • Experience correlating performance models or RTL/Verilator simulation against measured silicon or GPUs.
  • GPU kernel work in CUDA, CUTLASS, or Triton; familiarity with PyTorch internals.
  • Familiarity with PagedAttention, FlashAttention, speculative decoding, and disaggregated prefill.
  • Distributed inference experience (collectives, all-reduce, NCCL, NVLink, InfiniBand) or experience with MLPerf/production benchmarking pipelines.
  • Background at a hyperscaler, GPU vendor, accelerator company, or inference lab.

Education Requirements

BS or MS in Computer Engineering, Electrical Engineering, Computer Science, or equivalent practical experience.


About the Company

Company: Neurophos

Headquarters: Sunnyvale, California, United States

Startup developing silicon-photonic AI accelerator chips that use programmable metasurfaces and dense optical cells to perform matrix multiplications at the speed of light, targeting large-scale AI inference with much higher energy efficiency and performance than traditional electronic approaches.

Neurophos logo

Date Posted: 2026-09-04