Production Systems Engineer, NPI
Meta PlatformsJob Title
Production Systems Engineer, NPI
Role Summary
Join Meta's Release to Production (RTP) team to deliver end-to-end hardware lifecycle support for Meta servers, from prototyping and preproduction validation to production deployment and fleet health monitoring. The role collaborates with hardware and software teams, vendors, manufacturers, and operations to enable reliable compute and storage platforms in data centers.
Experience Level
Mid-level - typically requires 6+ years of relevant experience in hardware, systems, or storage domains.
Responsibilities
Work across engineering, vendors, and operations to validate, debug, and enable server platforms for production.
- Lead end-to-end system validation and deployment for compute and storage platforms (CPUs, SSD/HDD, DPUs) in data center environments.
- Create experiments and tooling to diagnose hardware, firmware, and software failures using telemetry, logs, and hands-on debugging.
- Design and implement system-level test plans (functional, stress, performance) including security-sensitive domains.
- Drive vendors and manufacturers to resolve hardware and firmware issues detected in production.
- Develop hardware fault management, error reporting, and handling strategies for server products.
- Analyze production failure data to identify trends and coordinate corrective actions.
- Create and maintain dashboards and visualizations to surface hardware health signals and remediation progress.
- Define test specifications and methodologies and improve test quality across internal teams and industry partners.
Requirements
Must-have technical skills and experience required to perform the role.
- 6+ years of relevant experience in domains such as ASIC development, compute (ARM/x86), AI/ML hardware, storage (HDD/SSD/HBA), memory (HBM/DDR5/LPDDR5), networking (NIC/DPU), or server interconnects.
- Experience with hardware fault management, error reporting, and error handling.
- Proficiency with Python, C/C++, or similar languages in a Linux environment for system management, automation, and CI/CD.
- Experience troubleshooting complex issues that cross hardware, firmware, and software boundaries.
- Ability to communicate technical findings to cross-functional teams and work effectively in a matrix organization.
Nice-to-have:
- Experience designing or validating HDD/SSD topologies and familiarity with the Linux block layer and storage RAS/security concepts.
- Working knowledge of bus protocols (I2C, SPI, USB, LPDDR, PCIe).
- Hands-on Linux server management and system-level debugging across multiple components.
- Experience authoring test plans for chipsets and integrating lab tools for automated workflows and large-scale deployments.
- Familiarity with SoC debugging tools (JTAG, GDB, DSTREAM, Trace32) and CI/CD tooling; scripting automation experience (2+ years preferred).
Education Requirements
Bachelor's degree in Computer Science, Computer Engineering, or a relevant technical field, or equivalent practical experience.
About the Company
Company: Meta Platforms
Headquarters: Menlo Park, California, United States
American technology company that develops social networking products (Facebook, Instagram, WhatsApp) and invests in virtual/augmented reality hardware and software through Reality Labs, focusing on connectivity, advertising, and immersive computing experiences.
