Senior Software Engineer, Network System Validation
NVIDIAJob Title
Senior Software Engineer, Network System Validation
Role Summary
Technical lead in Network System Validation for NVIDIA's Networking Business Unit (Mellanox). Own the validation roadmap, develop validation methodologies and automation frameworks, and lead hands-on debugging and performance analysis of networking technologies in large-scale AI cluster environments.
Work focuses on validating high-speed interconnects (Ethernet, InfiniBand, NVLink, DPUs) under real-world AI workloads.
Experience Level
Senior β 8+ years of relevant experience in networking, system validation, or related domains.
Responsibilities
Primary responsibilities include designing and executing validation plans, automation, and leading investigations across hardware and software stacks.
- Review requirements and design functional and performance validation plans for networking in large-scale AI clusters.
- Develop and maintain benchmarks, automation tools, and scripts for test execution, environment setup, log collection, and data analysis.
- Lead end-to-end investigation of complex issues: reproduce scenarios, analyze logs, telemetry, packet captures, and system metrics, and drive to root cause.
- Read and debug C/C++/Python source to investigate defects, validate fixes, and improve logging and instrumentation.
- Collaborate with hardware and software teams to debug NCCL, RoCE, RDMA, and related components.
- Profile AI training and inference workloads and correlate application behavior with network and system telemetry to identify performance and scalability limitations.
- Document findings and continuously improve validation methodologies, automation environments, and engineering processes.
- Mentor engineers and provide technical leadership on the validation roadmap.
Requirements
Must-have technical skills and experience; nice-to-have items listed below.
- 8+ years in networking, system validation, or related domains.
- Proven experience debugging complex production systems using hypothesis-driven experiments and driving issues to root cause.
- Ability to read, debug, and reason about C and C++ code (Rust or Go are a plus).
- Strong scripting and automation experience using Python, Bash, and/or Ansible.
- Deep understanding of distributed systems: concurrency, consistency models, fault tolerance, and large-scale performance under stress.
- Ability to drive technical alignment across teams and make high-quality architectural decisions quickly.
- Experience or interest in AI-driven test automation: intelligent scenario generation, LLM-augmented root-cause analysis, and autonomous validation pipelines.
Nice-to-have:
- Experience with large-scale clusters or distributed systems.
- Familiarity with NVIDIA networking products (ConnectX, SpecX, BlueField).
- Background in performance analysis, Kubernetes, or cloud environments.
- Experience with chaos testing, fault injection, or simulation systems.
Education Requirements
B.Sc. or B.A. in Computer Science, Electrical Engineering, or equivalent practical experience.
About the Company
Company: NVIDIA
Headquarters: Santa Clara, California, USA
NVIDIA is a global leader in accelerated computing, renowned for its innovative solutions in AI and digital twins that transform diverse industries. The company specializes in networking technologies, providing end-to-end InfiniBand and Ethernet solutions for servers and storage that optimize performance and scalability. NVIDIA serves sectors such as high-performance computing, enterprise data centers, and cloud computing, constantly reinventing its products and services to stay ahead in the market.
