AI Cluster Technical Program Manager – Validation, Debug & Agentic AI
Advanced Micro DevicesJob Title
AI Cluster Technical Program Manager – Validation, Debug & Agentic AI
Role Summary
Lead execution of AI cluster engineering programs focused on GPU platforms, rack-level solutions, and cluster validation. Coordinate hardware, firmware, networking, lab operations, and scale testing to ensure platform readiness for hyperscale and enterprise AI deployments.
Experience Level
Mid-level. The posting expects experience leading complex hardware or AI infrastructure programs but does not specify a required number of years.
Responsibilities
Manage end-to-end program delivery for GPU-based systems, rack bring-up, and cluster-scale validation while driving debug and incident response.
- Define, plan, and drive program schedules, dependency maps, resource forecasts, and risk/issue logs for server, rack, and cluster validation.
- Coordinate GPU, CPU, firmware, BIOS/BMC, system, networking, and lab teams for bring-up, EVT/DVT/PVT, and scale testing.
- Lead multi-node and multi-rack test planning, scheduling, coverage tracking, and readiness gates.
- Own rack-level delivery including compute trays, switch trays, cabling, power, cooling, and management infrastructure readiness.
- Drive failure analysis and system-level debug: triage, root-cause analysis, corrective actions, and closure tracking.
- Operate incident management for critical deployment and validation issues, run war rooms, and provide executive communications.
- Define and track operational metrics (MTTD, MTTM, MTTR, incident recurrence, fleet health) and lead post-incident reviews and preventive actions.
- Partner with scale, performance, and automation teams to ensure workloads, stress tests, and regression plans are ready for hardware arrival.
- Champion and coordinate Agentic AI / AIOps initiatives for automated triage, log analysis, root-cause identification, test orchestration, and incident management.
- Provide concise, data-driven status updates and drive cross-team escalations to remove blockers and mitigate risks.
Requirements
Must-have technical program management skills and domain knowledge to lead complex hardware and AI infrastructure programs.
- Proven experience leading complex hardware or AI infrastructure programs through bring-up, validation, and deployment phases.
- Strong technical understanding of GPU-based AI systems, rack architectures, and datacenter infrastructure.
- Demonstrated ability to drive debug execution, manage ambiguity, and lead cross-functional teams without direct authority.
- Strong written and verbal communication skills, including executive-level reporting.
- Proficiency with program management and execution tools (Jira, Confluence, dashboards, Excel, PowerPoint).
Nice-to-have:
- Hands-on experience with GPU cluster scale testing, system stress, or performance validation.
- Familiarity with rack-level bring-up, power/cooling constraints, networking, and failure modes at scale.
- Experience working through hardware/firmware debug cycles in pre-production or customer-facing environments.
- Experience managing fleet-scale validation and deployment for AI, cloud, HPC, or hyperscale infrastructure.
- Knowledge of infrastructure observability, telemetry pipelines, log analytics, and incident management frameworks.
Education Requirements
Bachelor's or Master's degree in systems, electrical engineering, computer science, or a related engineering discipline. PMP, Scrum Master, or equivalent program management training noted as desirable. (Equivalent practical experience likely acceptable based on general industry practice; the posting does not explicitly provide an equivalence statement.)
About the Company
Company: Advanced Micro Devices
Headquarters: Sunnyvale, California, USA
Advanced Micro Devices, or AMD, is a global semiconductor company that designs and manufactures microprocessors, graphics processors, and related technologies for a variety of computing devices. Known for pushing the boundaries of innovation, AMD's mission is to deliver high-performance computing solutions for AI, data centers, gaming, and embedded applications. They foster a collaborative, inclusive culture focused on creativity and problem-solving, aiming to drive progress and excellence in technology.
