Skip to main content
Meta Platforms logo

Production Systems Engineer, AI Systems

Meta Platforms
September 02, 2026
Full-time
On-site
Austin, Texas, United States
$144,000 - $204,000 USD yearly
Test Engineering Jobs, Level - Mid-Career

Job Title

Production Systems Engineer, AI Systems

Role Summary

Validate and scale next-generation AI and HPC server hardware from early bring-up through production readiness. Work with hardware design, firmware, software, networking, and capacity teams to drive system-level validation, root-cause analysis, and deployment acceptance for large-scale data center AI infrastructure.

Experience Level

Mid-level — typically requires 6+ years of relevant hardware systems engineering or system-level bring-up experience.

Responsibilities

Primary responsibilities center on end-to-end validation, debugging, and readiness of AI server platforms across hardware, firmware, and software layers.

  • Define and lead system validation strategies for AI/HPC platforms including accelerators, GPU clusters, and memory subsystems.
  • Perform hands-on bring-up, characterization, and validation of servers and components (PCIe, NVLink, DRAM, high-speed fabrics).
  • Create and maintain test specifications, validation procedures, and debug guides for NPI programs.
  • Investigate and root-cause complex failures spanning silicon, firmware, software, and hardware.
  • Track hardware and firmware defects through resolution while meeting NPI milestones.
  • Improve test coverage, tooling, and automation frameworks across the NPI lifecycle.
  • Partner with platform and capacity teams to define acceptance criteria and deployment readiness standards.
  • Drive data collection, analysis, and reporting to identify systemic quality trends and inform go/no-go decisions.
  • Communicate validation status, risks, and technical findings to engineering teams and vendors.
  • Collaborate on hardware–software interface requirements for telemetry, diagnostics, and remote management.

Requirements

Must-have technical experience and skills required for the role; preferred items listed separately.

  • 6+ years experience in hardware systems engineering, silicon or firmware validation, or system-level bring-up for AI servers, GPUs, TPUs, or AI accelerators.
  • Experience in one or more: ASIC bring-up/characterization, board-level debug, firmware validation, or large-scale system validation in data centers.
  • Proven experience developing test specifications, validation procedures, and debug methodologies for complex hardware systems.
  • Experience leading root-cause analysis and troubleshooting across hardware, firmware, and software stacks.
  • Hands-on knowledge of high-speed interconnects or memory subsystems (PCIe, NVLink, DDR5, HBM) in AI/HPC contexts.
  • Experience analyzing system telemetry and fleet health data to identify reliability trends and drive improvements.
  • Nice-to-have: proficiency in Python for automation and data analysis.
  • Nice-to-have: familiarity with Linux-based servers, data center management tooling, and remote diagnostics/telemetry interfaces.
  • Nice-to-have: experience with InfiniBand or similar high-speed fabrics in production environments.

Education Requirements

Bachelor's degree in Computer Science, Computer Engineering, or a relevant technical field — or equivalent practical experience. (No additional degree or certification requirements specified.)


About the Company

Company: Meta Platforms

Headquarters: Menlo Park, California, United States

American technology company that develops social networking products (Facebook, Instagram, WhatsApp) and invests in virtual/augmented reality hardware and software through Reality Labs, focusing on connectivity, advertising, and immersive computing experiences.

Meta Platforms logo

Date Posted: 2026-09-02