Production Systems Engineer, AI Systems
Meta PlatformsJob Title
Production Systems Engineer, AI Systems
Role Summary
Validate and scale next-generation AI and HPC server hardware from early bring-up through production readiness. Work with hardware design, firmware, software, networking, and capacity teams to drive system-level validation, root-cause analysis, and deployment acceptance for large-scale data center AI infrastructure.
Experience Level
Mid-level — typically requires 6+ years of relevant hardware systems engineering or system-level bring-up experience.
Responsibilities
Primary responsibilities center on end-to-end validation, debugging, and readiness of AI server platforms across hardware, firmware, and software layers.
- Define and lead system validation strategies for AI/HPC platforms including accelerators, GPU clusters, and memory subsystems.
- Perform hands-on bring-up, characterization, and validation of servers and components (PCIe, NVLink, DRAM, high-speed fabrics).
- Create and maintain test specifications, validation procedures, and debug guides for NPI programs.
- Investigate and root-cause complex failures spanning silicon, firmware, software, and hardware.
- Track hardware and firmware defects through resolution while meeting NPI milestones.
- Improve test coverage, tooling, and automation frameworks across the NPI lifecycle.
- Partner with platform and capacity teams to define acceptance criteria and deployment readiness standards.
- Drive data collection, analysis, and reporting to identify systemic quality trends and inform go/no-go decisions.
- Communicate validation status, risks, and technical findings to engineering teams and vendors.
- Collaborate on hardware–software interface requirements for telemetry, diagnostics, and remote management.
Requirements
Must-have technical experience and skills required for the role; preferred items listed separately.
- 6+ years experience in hardware systems engineering, silicon or firmware validation, or system-level bring-up for AI servers, GPUs, TPUs, or AI accelerators.
- Experience in one or more: ASIC bring-up/characterization, board-level debug, firmware validation, or large-scale system validation in data centers.
- Proven experience developing test specifications, validation procedures, and debug methodologies for complex hardware systems.
- Experience leading root-cause analysis and troubleshooting across hardware, firmware, and software stacks.
- Hands-on knowledge of high-speed interconnects or memory subsystems (PCIe, NVLink, DDR5, HBM) in AI/HPC contexts.
- Experience analyzing system telemetry and fleet health data to identify reliability trends and drive improvements.
- Nice-to-have: proficiency in Python for automation and data analysis.
- Nice-to-have: familiarity with Linux-based servers, data center management tooling, and remote diagnostics/telemetry interfaces.
- Nice-to-have: experience with InfiniBand or similar high-speed fabrics in production environments.
Education Requirements
Bachelor's degree in Computer Science, Computer Engineering, or a relevant technical field — or equivalent practical experience. (No additional degree or certification requirements specified.)
About the Company
Company: Meta Platforms
Headquarters: Menlo Park, California, United States
American technology company that develops social networking products (Facebook, Instagram, WhatsApp) and invests in virtual/augmented reality hardware and software through Reality Labs, focusing on connectivity, advertising, and immersive computing experiences.
