Skip to main content
NVIDIA logo

Senior Software Engineer, AI Inference Systems

NVIDIA
September 19, 2026
Full-time
Remote
Worldwide
292,500 zł - 650,000 zł PLN yearly
EDA Jobs, Level - Senior

Job Title

Senior Software Engineer, AI Inference Systems

Role Summary

Join the team building high-performance AI inference systems that serve large-scale models efficiently. The role focuses on architecting and implementing inference stacks, optimizing GPU kernels and compilers, and scaling workloads across multi-GPU, multi-node, and multi-cloud environments.

You will work with inference, compiler, scheduling, and performance teams to improve throughput, latency, and resource utilization for production ML workloads.

Experience Level

Senior. Typical guidance: senior-level engineers with ~7+ years of relevant experience; candidates with a Master’s and ~5+ years or a PhD with research record are also relevant.

Responsibilities

Key responsibilities include designing, implementing, and optimizing inference software and infrastructure.

  • Implement and extend features in inference frameworks (e.g., vLLM); profile and optimize frameworks for latency and throughput using methods like speculative decoding, tensor/expert/pipeline parallelism, and prefill-decode disaggregation.
  • Develop, optimize, and benchmark GPU kernels (hand-tuned and compiler-generated) using fusion, autotuning, and memory/layout optimizations; improve developer-facing DSLs and compiler infrastructure to approach peak hardware utilization.
  • Define and build benchmarking methodologies and tooling; contribute benchmarks and submissions to industry suites such as MLPerf Inference.
  • Architect scheduling and orchestration for containerized large-scale inference deployments across GPU clusters and clouds.
  • Conduct and publish original research; evaluate recent publications and integrate promising research prototypes into products.
  • Collaborate across teams (inference, compiler, scheduling, performance) to scale workloads across multi-GPU, multi-node, and multi-cloud environments.

Requirements

Must-have technical skills and experience; a short list of differentiators follows.

  • Strong programming skills in Python and C/C++; experience with Go or Rust is a plus.
  • Solid CS fundamentals: algorithms & data structures, operating systems, computer architecture, parallel programming, and distributed systems.
  • Performance engineering experience with ML frameworks and inference engines (e.g., PyTorch, vLLM).
  • GPU programming and performance experience: CUDA, memory hierarchy, streams, NCCL; familiarity with profiling/debugging tools (e.g., Nsight Systems/Compute).
  • Experience with containers and orchestration (Docker, Kubernetes, Slurm) and familiarity with Linux namespaces and cgroups.
  • Proven debugging, problem-solving, and communication skills in fast-paced, cross-functional settings.

Nice-to-have

  • Hands-on experience building and optimizing LLM inference engines (e.g., vLLM, SGLang).
  • Experience with ML compilers and DSLs (e.g., Triton, TorchDynamo/Inductor, MLIR/LLVM, XLA) and GPU libraries/features (CUTLASS, CUDA Graph, Tensor Cores).
  • Experience with containerization/virtualization internals (containerd/CRI-O/CRIU).
  • Experience with cloud platforms (AWS/GCP/Azure), infrastructure as code, CI/CD, and production observability.
  • Contributions to open-source projects or relevant publications.

Education Requirements

Bachelor's, Master's or PhD in Computer Science, Computer Engineering, Software Engineering, or a closely related technical field; or equivalent practical experience. PhD candidates should have relevant thesis work and publications in ML systems, GPU architecture, or high-performance computing.


About the Company

Company: NVIDIA

Headquarters: Santa Clara, California, USA

NVIDIA is a global leader in accelerated computing, renowned for its innovative solutions in AI and digital twins that transform diverse industries. The company specializes in networking technologies, providing end-to-end InfiniBand and Ethernet solutions for servers and storage that optimize performance and scalability. NVIDIA serves sectors such as high-performance computing, enterprise data centers, and cloud computing, constantly reinventing its products and services to stay ahead in the market.

NVIDIA logo

Date Posted: 2026-09-18