Systems Performance Modeling Engineer
TensordyneJob Title
Systems Performance Modeling Engineer
Role Summary
Build and maintain simulation and analytical models that predict generative AI inference performance from a single accelerator to rack, pod, and cluster scale. Work closely with architects and hardware, networking, and software teams to capture workload behavior, validate models against real hardware, and provide data-driven guidance for design and configuration decisions.
Experience Level
Mid-level - expects several years of hands-on experience building performance models, simulators, or analytical tools for ML workloads, distributed systems, or computer architecture. No explicit years listed in the posting.
Responsibilities
Primary responsibilities include modeling, experimentation, validation, and documentation of system-level performance for generative AI inference.
- Implement and extend simulation-based models for compute, memory, collective communication, and network fabric at rack/pod/cluster scale.
- Model interactions of serving strategies (tensor/pipeline/expert parallelism, prefill/decode disaggregation, batching, KV-cache placement) with silicon and fabric topology.
- Build trace-capture and replay tools to record runtime execution and replay under hypothetical system configurations.
- Model and implement collective communication algorithms for multi-hop scale-out fabrics.
- Run calibration experiments on hardware, compare measurements to model predictions, and resolve discrepancies.
- Perform design-space sweeps and produce clear analyses to inform ASIC, fabric, and system configuration decisions.
- Maintain a fast, tested, and reproducible modeling codebase that other engineers can run.
Requirements
Must-have technical skills and abilities required to perform the role.
- Hands-on experience building performance models, simulators, or analytical tools for ML workloads, distributed systems, or computer architecture.
- Solid understanding of distributed ML execution and parallelism strategies, including collective communication (All-Reduce, All-Gather, All-to-All) and scalability behavior.
- Working knowledge of system architecture across compute, memory, interconnect, and networking, and ability to reason about cross-layer bottlenecks.
- Experience comparing model predictions to real measurements and debugging divergences.
- Strong programming skills in C++ and Python with emphasis on clean, testable, maintainable code.
- Ability to design experiments from loosely defined questions and deliver results with minimal supervision.
- Clear written and verbal communication for presenting data and trade-offs to engineering audiences.
Nice-to-have:
- Familiarity with LLM inference serving (batching, KV-cache management, disaggregated prefill/decode).
- Experience modeling or benchmarking collective communication libraries (NCCL, RCCL, or similar) on real clusters.
- Background in data center or HPC networking: topologies, RDMA/RoCE, and congestion behavior.
- Experience profiling ML workloads on accelerators (GPUs, TPUs, or custom ASICs).
- Exposure to hardware/software co-design, early-stage architecture evaluation, or relevant publications/open-source contributions.
Education Requirements
MS or higher in Computer Science, Computer Engineering, Electrical Engineering, or a related field (as stated in the posting).
About the Company
Company: Tensordyne
Tensordyne is a startup developing high-performance, low-power AI inference processors and multi-chip systems for generative AI acceleration in data centers. The company focuses on ASIC design, verification, and integration of computational accelerators with third-party SoC IP.
