Skip to main content
NVIDIA logo

Senior Software Engineer - Distributed Systems Engineer, EDA Infrastructure

NVIDIA
September 01, 2026
Full-time
Remote
Worldwide
EDA Jobs, Level - Senior

Job Title

Senior Software Engineer - Distributed Systems Engineer, EDA Infrastructure

Role Summary

Design, build, and operate scalable automation and platform services that manage large fleets of GPU- and CPU-based compute infrastructure for Electronic Design Automation (EDA) workloads.

Work across software, operating systems, schedulers, networking, storage, and hardware teams to improve reliability, availability, and operational efficiency of production compute environments.

Experience Level

Senior — expected to have substantial independent experience; role targets candidates with strong infrastructure or software engineering backgrounds. (See Education section for degree information.)

Responsibilities

Primary responsibilities include designing and operating systems that automate lifecycle and recovery for large-scale compute infrastructure used for chip design workloads.

  • Design and build platforms to automate provisioning, configuration, operation, and lifecycle management of GPU and CPU compute infrastructure.
  • Develop monitoring, health-management, and automated remediation systems to improve reliability, availability, and utilization.
  • Automate hardware deployment, OS configuration, firmware and software updates, cluster enrollment, and recovery workflows.
  • Integrate services and workflows with workload schedulers, infrastructure management systems, and observability platforms.
  • Use diagnostics, OS signals, scheduler data, and network/storage telemetry to identify failures and return systems to service.
  • Participate in incident response, root-cause analysis, capacity planning, and continuous improvement of production services.
  • Collaborate with EDA, infrastructure, networking, storage, and hardware engineering teams to deliver scalable solutions for critical workloads.

Requirements

Must-have skills and experience to perform the role effectively; followed by relevant nice-to-have qualifications.

  • Must-have: Strong programming experience in Go or Python, with sound knowledge of data structures, algorithms, testing, and software design.
  • Experience designing automation for distributed systems and managing large fleets of Linux-based compute nodes.
  • Knowledge of performance, security, reliability, fault tolerance, state management, and data consistency in complex systems.
  • Experience with infrastructure automation, software deployment, observability, and operational recovery.
  • Strong communication skills and ability to work effectively across teams and geographic regions.
  • Systematic problem-solving approach, strong sense of ownership, and focus on reducing operational toil.
  • Nice-to-have: Experience designing or operating large-scale EDA or HPC infrastructure; deep knowledge of Linux, GPU/CPU server architecture, networking, storage, and bare-metal lifecycle management.
  • Hands-on experience with workload schedulers and cluster-management platforms such as Slurm, LSF, Kubernetes, or Bright Cluster Manager.
  • Experience with EDA applications, license-management systems, high-throughput batch workloads, or semiconductor design workflows.
  • Experience building automated health checks, remediation workflows, firmware and OS upgrade systems, or node-provisioning systems across multiple data centers or heterogeneous hardware.

Education Requirements

BS in Computer Science, Engineering, Physics, Mathematics, or a related field; or equivalent practical experience.


About the Company

Company: NVIDIA

Headquarters: Santa Clara, California, USA

NVIDIA is a global leader in accelerated computing, renowned for its innovative solutions in AI and digital twins that transform diverse industries. The company specializes in networking technologies, providing end-to-end InfiniBand and Ethernet solutions for servers and storage that optimize performance and scalability. NVIDIA serves sectors such as high-performance computing, enterprise data centers, and cloud computing, constantly reinventing its products and services to stay ahead in the market.

NVIDIA logo

Date Posted: 2026-08-31