Member of Technical Staff, Performance Modeling
NetpremeJob Title
Member of Technical Staff, Performance Modeling
Role Summary
Develop functional and performance models for a scale-up network-attached memory expansion device for AI accelerators. Work on the silicon architecture team to explore design tradeoffs, validate performance assumptions, and identify bottlenecks across compute, memory, and interconnects.
This role is performed onsite at one of the company offices in Santa Clara, CA or Boston, MA.
Experience Level
Mid-level; typically 5–10+ years of relevant experience in performance modeling for data-movement devices.
Responsibilities
Primary responsibilities include building and maintaining performance models and collaborating with architecture, system, and workload teams.
- Build and maintain system- and chip-level performance models for high-bandwidth data movement in scale-up environments.
- Model workloads from software memory access patterns through network distribution to on-device memory channels.
- Collaborate daily with silicon architects, system designers, and workload owners to align performance expectations and constraints.
- Identify performance bottlenecks, scaling limits, and sensitivity points across compute, memory, and interconnects in end-to-end workload settings.
- Document and clearly communicate modeling assumptions, limitations, and conclusions to technical and non-specialist stakeholders.
Requirements
Must-have skills and experience; preferred items listed separately.
Must-have
- 5–10+ years of experience in performance modeling for data-movement devices (NICs, memory expansion cards such as CXL, IPU/DPU, NoC).
- Ability to reason across multiple abstraction layers, from architectural details to system-level performance behavior.
- Experience exploring design tradeoffs and validating performance assumptions early in the development cycle.
- Strong analytical and communication skills to present modeling assumptions and results.
Nice-to-have
- Experience modeling networking protocols with memory semantics.
- Familiarity with shared memory systems and frameworks (e.g., CUDA VMM).
- Familiarity with modern AI/ML workload behavior (LLM inference and sharding, KV caching, serving systems).
- Experience with scale-up high-bandwidth interconnects (e.g., NVLink) and modeling memory subsystems.
Education Requirements
Bachelor’s or Master’s degree in Electrical Engineering, Computer Engineering, or a closely related field.
About the Company
Company: Netpreme
Headquarters: Santa Clara, CA, United States
Early-stage startup developing ASIC memory-acceleration solutions for AI and data-center workloads. Focused on designing high-performance silicon, IP and subsystems to improve efficiency and performance for large language models and other AI applications.
