Job Title
System Software Engineer, Distributed Systems
Role Summary
Join the VLSI Productivity and Infrastructure team to design and operate userspace system software that supports large-scale chip design workflows on bare-metal Linux hosts. The role focuses on distributed systems, state coordination over shared filesystems, orchestration around batch schedulers, reliability, performance, and incremental modernization of legacy codebases.
Experience Level
Mid-level. The posting requests approximately 5+ years of relevant experience.
Responsibilities
Deliver and operate core userspace infrastructure for long-running engineering workflows and enable 1000+ chip designers to be more productive.
- Design, build, and maintain core components of productivity platforms and workflow infrastructure.
- Develop reliable userspace services for long-running workflows on bare-metal Linux (no privileged ops or containers).
- Implement state coordination over NFS: atomicity, idempotency/dedup, and partial-write recovery without privileged operations.
- Build and improve orchestration around IBM LSF: submission/tracking, retries/cancel, log capture, fairness, and backpressure.
- Perform incremental modernization of legacy codebases (e.g., migrating Perl to Go) with staged rollouts and observability.
- Debug and improve performance and reliability across Linux and Kubernetes; build operational tooling and observability.
- Work with engineering users to turn ambiguous workflows into durable, measurable production systems.
Requirements
Must-have technical skills and experience.
- 5+ years developing and operating production software in Go and/or Python, ideally in large codebases.
- Strong Linux fundamentals: processes, filesystems, permissions, synchronization/locks, concurrency, and debugging.
- Solid distributed-systems thinking: handling failures, retries/timeouts, backoff, idempotency, and operational rigor.
- Experience building long-runtime automation or services on shared compute clusters (batch schedulers, build systems).
- Ability to plan safe deliveries: instrumentation, staged rollout strategies, parity checks, and measurable outcomes.
Nice-to-have:
- Hands-on experience with shared filesystems at scale (NFS) or coordination on eventually-consistent storage.
- Experience with batch job scheduling, shared compute fleets, or large-scale build systems.
- Track record of incremental modernization: tests, shadow runs, canaries, and rollback plans.
- Experience optimizing metadata-heavy systems and reducing I/O or R/W hot spots.
- Strong incident response and debugging: root-cause analysis, remediation, and rapid ownership of unfamiliar codebases.
Education Requirements
B.S. in Computer Science or Electrical Engineering preferred, or equivalent practical experience. (The posting explicitly allows equivalent experience.)
About the Company
Company: NVIDIA
Headquarters: Santa Clara, California, USA
NVIDIA is a global leader in accelerated computing, renowned for its innovative solutions in AI and digital twins that transform diverse industries. The company specializes in networking technologies, providing end-to-end InfiniBand and Ethernet solutions for servers and storage that optimize performance and scalability. NVIDIA serves sectors such as high-performance computing, enterprise data centers, and cloud computing, constantly reinventing its products and services to stay ahead in the market.

Date Posted: 2026-08-22