NVIDIA logo

Senior Compute Platform Engineer, LSF

NVIDIA
August 22, 2026
Full-time
Remote friendly (Santa Clara, California, United States)
Worldwide
$184,000 - $356,500 USD yearly
Other Semiconductor Jobs, Level - Senior

Job Title

Senior Compute Platform Engineer, LSF

Role Summary

Deep-specialist engineering role owning scheduler behavior and performance for NVIDIA's federated LSF compute farm. You will diagnose and resolve scheduling latency and contention across multiple LSF cells, set technical designs for cell topology and federation, and work with infrastructure, IaC, and CAD teams to ensure reliable large-scale batch compute for EDA workloads.

Applications accepted through 2026-08-24.

Experience Level

Senior-level. Typical requirement: 8+ years in HPC or large-scale batch compute, including 5+ years using IBM Spectrum LSF.

Responsibilities

Primary responsibilities include ownership, diagnosis, design, and collaboration to maintain and scale the LSF estate:

  • Own scheduler behavior across 15–25 federated LSF cells; tune mbatchd and mbschd and analyze scheduling cycles.
  • Diagnose MultiCluster forwarding issues, remote queue sizing, forwarding policy, and cross-cluster pend behavior.
  • Analyze contention patterns as cells approach host-count ceilings and identify root causes of latency.
  • Define cell topology and federation strategy; decide when to split or merge cells.
  • Collaborate with IaC engineers to encode scheduler policy into robust configuration schemas for MultiCluster scale.
  • Work with CAD and methodology teams on workloads that break assumptions: very large memory jobs, interactive vs batch contention, and tape-out bursts.
  • Serve as the escalation point for complex scheduler and performance incidents.

Requirements

Must-have technical skills and experience:

  • 8+ years in HPC or large-scale batch compute, with 5+ years operating IBM Spectrum LSF in production.
  • Demonstrated depth in LSF internals; able to explain and debug a scheduling cycle from submission to dispatch.
  • Hands-on MultiCluster experience in production, multi-site environments.
  • Strong Linux systems fundamentals and system programming experience.
  • Proficient scripting in Python, Perl, and shell.
  • Excellent problem-solving and incident troubleshooting under production load.

Nice-to-have:

  • Experience developing LSF or working in escalation engineering rather than only as a user.
  • Familiarity with LSF integration points: esub, eexec, elim, submit wrappers, RTM, or LSF APIs.
  • Background in semiconductor or EDA compute environments and experience migrating estates from Slurm, PBS, or Grid Engine without user-visible outages.

Education Requirements

Bachelor's or Master's degree in Computer Science, Computer Engineering, or a related field, or equivalent practical experience.


About the Company

Company: NVIDIA

Headquarters: Santa Clara, California, USA

NVIDIA is a global leader in accelerated computing, renowned for its innovative solutions in AI and digital twins that transform diverse industries. The company specializes in networking technologies, providing end-to-end InfiniBand and Ethernet solutions for servers and storage that optimize performance and scalability. NVIDIA serves sectors such as high-performance computing, enterprise data centers, and cloud computing, constantly reinventing its products and services to stay ahead in the market.

NVIDIA logo

Date Posted: 2026-08-22