System Software Engineer, Node & Cluster Management
MatXJob Title
System Software Engineer, Node & Cluster Management
Role Summary
Work on host system software to make MatX AI systems manageable at node and cluster scale. The team owns Linux userspace, management daemons, node-level APIs, and co-designs management behavior with BMC/OpenBMC firmware engineers.
This role focuses on building node-level management, cluster management, operator tooling, and lab automation during hardware bring-up. Candidates must be authorized to work in the United States and work from our Mountain View office Tuesdays–Thursdays.
Experience Level
Senior-level. The posting requests 8+ years of systems software experience and deep low-level systems expertise.
Responsibilities
Deliver node and cluster management features, tooling, and automation for production and bring-up environments.
- Design and implement node-level management plane: expose health, inventory, telemetry, and controls via HTTP/REST endpoints (Redfish-style or custom).
- Design cluster management and failover algorithms to minimize downtime and support fleet-wide aggregation and alerting.
- Build CLI utilities and operator tools for diagnostics, firmware updates, recovery, and daily operations.
- Partner with BMC firmware engineers to present unified in-band/out-of-band management and orchestrate firmware updates and recovery flows.
- Work below the API layer as needed: telemetry daemons, driver interfaces, raw device access to prototype, debug, and unblock development.
- Develop tooling and automation for lab provisioning, test orchestration, and regression monitoring during hardware bring-up.
- Define software contracts between on-node daemons, BMC stack, and management services.
- Debug production issues spanning APIs, daemons, kernel drivers, firmware, and hardware.
Requirements
Must-have technical skills and experience required to perform the role.
- 8+ years of systems software experience (low-level Linux-focused work required).
- Hands-on Linux systems development and userspace experience; comfortable reading and debugging kernel driver and daemon code.
- Strong programming in C and at least one systems/service language: Go, Rust, C++, or Python.
- Experience designing and building HTTP/REST APIs and CLI tools for hardware or infrastructure management.
- Proven debugging across API, daemon, kernel, firmware, and hardware boundaries.
- Ability to collaborate with firmware engineers and align host-side and BMC-side management capabilities.
- Self-driven and pragmatic: able to implement working management endpoints against new hardware with minimal specification.
Nice-to-have:
- Experience with Redfish, OpenBMC, gNMI, IPMI, or other datacenter hardware management standards.
- Cluster/fleet management experience for GPU or accelerator infrastructure.
- Hardware bring-up, lab automation, or manufacturing/qualification test infrastructure experience.
- Familiarity with firmware update orchestration, secure boot, or attestation flows.
Education Requirements
BS or higher in Computer Science, Electrical Engineering, or equivalent practical experience. (The posting explicitly allows equivalent practical experience in lieu of a degree.)
About the Company
Company: MatX
Headquarters: Mountain View, California, USA
MatX specializes in creating faster chips for large language models (LLMs), focusing on innovative hardware and software solutions. The company fosters a collaborative and supportive work environment, welcoming candidates of all experience levels. Their approach prioritizes deep understanding and consideration of novel methods to drive efficiency and performance in their projects, particularly in silicon design and related engineering roles.
