NVIDIA logo

Manager, Infrastructure Engineering and DevOps

NVIDIA
August 19, 2026
Full-time
On-site
Yokne'am Illit, Israel
Other Semiconductor Jobs, Level - Senior

Job Title

Manager, Infrastructure Engineering and DevOps

Role Summary

Lead an infrastructure engineering team that builds and operates bare-metal and VM provisioning, CI/CD, server fleet automation, and high-performance networking environments to support firmware, driver, hardware, software, and verification R&D teams.

The role combines technical leadership, hands-on problem resolution, roadmap and delivery ownership, and people management to ensure reliable infrastructure at scale.

Experience Level

Senior-level manager. The role expects ~8+ years of relevant engineering experience and 3+ years of people or technical project leadership.

Responsibilities

Own the technical direction, delivery, and operational support of infrastructure platforms used by internal engineering teams.

  • Lead, mentor, and grow a team responsible for provisioning (bare-metal and VMs), server fleet automation, CI/CD, and lab/network environments.
  • Define and execute the team roadmap, set priorities, and manage delivery commitments across multiple initiatives.
  • Manage VM and image lifecycle across multiple Linux distributions: image readiness, OS compatibility, package baselines, kernel and boot configurations.
  • Build and maintain infrastructure to enable efficient provisioning, testing, validation, and debug workflows for engineering teams.
  • Provide customer-facing support: triage, root-cause analysis, bottleneck removal, and workflow optimization with clear stakeholder communication.
  • Lead complex system debug and recovery for server bring-up, firmware/driver interactions, boot/network failures, and lab instability.
  • Ensure production readiness: inventory management, resource allocation, observability, and automated recovery practices.
  • Partner with firmware, driver, hardware, software, cloud, and verification teams to capture requirements and improve reliability.

Requirements

Key qualifications and skills required for success in this role.

  • Must-have: 8+ years in Linux systems administration, infrastructure automation, DevOps, system software, firmware infrastructure, lab infrastructure, or related domains.
  • Must-have: 3+ years leading or managing engineering teams, technical projects, or cross-functional infrastructure initiatives.
  • Must-have: Strong Linux technical skills: systemd, package management, kernel parameters, GRUB, sysctl tuning, NFS, networking, boot flows, and service management.
  • Must-have: Hands-on experience building and debugging automation using Python and scripting, and working with CI/CD workflows and modern development practices.
  • Must-have: Experience managing infrastructure across multiple Linux distributions: OS image management, provisioning flows, package dependency resolution, and environment consistency.
  • Must-have: Proven ability to support internal customers: incident triage, root-cause analysis, flow optimization, and cross-team communication.
  • Must-have: Strong people leadership: coaching, performance management, hiring, prioritization, and building inclusive teams.
  • Nice-to-have: Experience with RDMA/InfiniBand, OFED, SR-IOV, PCI passthrough, VFIO/IOMMU, or other high-speed networking technologies.
  • Nice-to-have: Familiarity with Ansible, infrastructure-as-code, Jenkins, Kubernetes, Docker, KVM/QEMU, libvirt, Vagrant, or multi-architecture environments (x86_64, aarch64, ppc64le).
  • Nice-to-have: Experience with firmware R&D, hardware bring-up, driver development, lab automation, Redfish/IPMI/iDRAC/iLO, BMC/BIOS automation, remote power control, or fleet recovery workflows.
  • Nice-to-have: Prior exposure to NVIDIA/Mellanox hardware and tooling (e.g., ConnectX, BlueField) is beneficial but not required.

Education Requirements

B.Sc. in Computer Engineering, Computer Science, Electrical Engineering, or a related technical field, or equivalent practical experience.


About the Company

Company: NVIDIA

Headquarters: Santa Clara, California, USA

NVIDIA is a global leader in accelerated computing, renowned for its innovative solutions in AI and digital twins that transform diverse industries. The company specializes in networking technologies, providing end-to-end InfiniBand and Ethernet solutions for servers and storage that optimize performance and scalability. NVIDIA serves sectors such as high-performance computing, enterprise data centers, and cloud computing, constantly reinventing its products and services to stay ahead in the market.

NVIDIA logo

Date Posted: 2026-08-19