Job Title
Senior Software and System Architect
Role Summary
Define software and system architecture for NVIDIA datacenter DPU/NIC platforms with emphasis on DPU management, QoS, performance, telemetry, and observability. Collaborate with hardware, firmware, driver, system engineering, validation, product management and customers from early architecture through pre-silicon design, bring-up, and production readiness.
Experience Level
Senior β typically requires 9+ years of experience in networking, system or embedded software, firmware, or datacenter infrastructure.
Responsibilities
Primary responsibilities include architecture definition, specification, and cross-team leadership:
- Own software and system architecture for next-generation DPU management, QoS, performance, telemetry, and observability features.
- Define end-to-end control and management flows across DOCA, host drivers, embedded firmware, BMC, management controllers and external management systems.
- Specify telemetry and observability requirements: counters, logs, traces, events, health monitoring, debug data, and streaming telemetry.
- Define management interfaces and APIs for configuration, provisioning, lifecycle operations, diagnostics, and field serviceability.
- Produce clear architecture specifications, interface definitions, flow diagrams, and design documents for software, firmware, and system teams.
- Partner with R&D and cross-functional teams to translate architecture into implementable designs and guide features through development, validation, silicon bring-up, and production.
- Analyze system performance bottlenecks, interoperability issues, telemetry gaps, and customer-reported problems; incorporate findings into future architecture.
- Coordinate with system and cluster architects to ensure NIC/DPU features integrate into end-to-end datacenter networking designs.
Requirements
Must-have and preferred qualifications (concise):
-
Must-have: 9+ years experience in networking, system software, embedded software, firmware, or datacenter infrastructure.
-
Must-have: Proven software architecture or technical leadership experience and strong written/verbal communication skills.
-
Must-have: Deep understanding of networking concepts and protocols such as Ethernet, TCP/IP, RDMA/RoCE, congestion control, QoS, virtualization overlays, and traffic management.
-
Must-have: Strong background with DPUs, SmartNICs, or other high-performance networking devices.
-
Must-have: Experience with system management, provisioning, monitoring, telemetry, diagnostics, and lifecycle-management flows; familiarity with Redfish, PLDM, MCTP, IPMI, gNMI, SNMP, Netconf, REST, or gRPC-based APIs.
-
Nice-to-have: Hands-on Linux networking, device drivers, embedded Linux, BMC software, DOCA, DPDK, OVS, Kubernetes networking; performance profiling, eBPF, Prometheus/Grafana experience.
-
Nice-to-have: Experience with performance counters, large-scale telemetry systems, RAS/diagnosability/serviceability, or GAI-based telemetry analysis tools.
Education Requirements
B.Sc. or M.Sc. in Computer Engineering, Computer Science, Electrical Engineering, or equivalent practical experience.
About the Company
Company: NVIDIA
Headquarters: Santa Clara, California, USA
NVIDIA is a global leader in accelerated computing, renowned for its innovative solutions in AI and digital twins that transform diverse industries. The company specializes in networking technologies, providing end-to-end InfiniBand and Ethernet solutions for servers and storage that optimize performance and scalability. NVIDIA serves sectors such as high-performance computing, enterprise data centers, and cloud computing, constantly reinventing its products and services to stay ahead in the market.

Date Posted: 2026-07-31