Senior Solutions Architect, CSP System
NVIDIAJob Title
Senior Solutions Architect, CSP System
Role Summary
Senior technical lead on NVIDIA's Cloud Service Provider (CSP) Solutions Architect team in China, focused on GPU and AI infrastructure for hyperscale workloads. The role drives system-level solution optimization, technical strategy, and high-value engagement with major Chinese CSPs to accelerate large-scale AI training, inference, Agentic AI, gaming AI and distributed computing deployments.
Experience Level
Senior β typically 7+ years of relevant hands-on experience in GPU architecture, AI systems, large-scale data center infrastructure, or hyperscale cloud computing.
Responsibilities
Provide technical leadership and customer-facing solutions across GPU/AI infra for Chinese CSPs. Key responsibilities include:
- Act as primary technical authority for NVIDIA GPU systems and full-stack AI infrastructure for top-tier Chinese CSP accounts.
- Design and optimize GPU cluster architectures, heterogeneous compute configurations, and end-to-end AI workload pipelines.
- Perform workload bottleneck analysis and implement system-, kernel-, and framework-level tuning for training, inference, RL and gaming workloads.
- Enable CPU+GPU co-optimization (Vera/Grace) to reduce data-movement bottlenecks and improve RL/Agentic AI throughput and latency.
- Lead technical workshops, hands-on trainings, PoCs and production pilots; produce reference designs and deployment guidelines for mass rollout.
- Collaborate with sales, BD and product teams to drive technical penetration and translate customer requirements into product roadmap input.
- Contribute upstream patches and best practices to open-source AI infra projects and develop China-localized operational standards.
- Mentor junior SAs and standardize CSP engagement and solution delivery processes.
Requirements
Must-have technical skills and experience; items in parentheses are nice-to-have where noted.
- 7+ years of hands-on experience with GPU architecture, AI system optimization, large-scale data center infrastructure, or hyperscale cloud computing.
- Deep knowledge of GPU microarchitecture, CUDA programming model, memory hierarchy, and system scheduling; strong performance profiling and bottleneck analysis skills.
- Proficient in C/C++ and Python; experience with CUDA kernels, compiler toolchains, and AI framework optimization (PyTorch, TensorRT).
- Experience tuning distributed systems at scale, including NCCL and large-scale training/inference frameworks.
- Proven track record working with major CSPs or hyperscalers and understanding of public cloud AI service architectures and cluster operations.
- Excellent technical communication and presentation skills for both engineering and business stakeholders.
- Strong cross-functional collaboration and project ownership; hands-on engineering capability to drive end-to-end technical deliveries.
- (Nice-to-have) Familiarity with NVIDIA full-stack products (GPU data center hardware, TensorRT-LLM, Dynamo, NCCL, CUDA) and cluster networking diagnostics.
- (Nice-to-have) Experience with Linux kernel, drivers, virtualization for containerized fleets, or open-source contributions to AI infra projects.
Education Requirements
Bachelor's, Master's, or PhD in Computer Science, Computer Engineering, Electrical Engineering, or a related field β or equivalent industry experience.
About the Company
Company: NVIDIA
Headquarters: Santa Clara, California, USA
NVIDIA is a global leader in accelerated computing, renowned for its innovative solutions in AI and digital twins that transform diverse industries. The company specializes in networking technologies, providing end-to-end InfiniBand and Ethernet solutions for servers and storage that optimize performance and scalability. NVIDIA serves sectors such as high-performance computing, enterprise data centers, and cloud computing, constantly reinventing its products and services to stay ahead in the market.
