Job Title
HPC and AI System Manager – Field Application Engineering
Role Summary
Operate and maintain AMD HPC/AI systems used by Field Application Engineering, customers, and OEM partners. The role focuses on regional hands-on support while contributing to global, follow-the-sun operations and automation.
Work with engineering, architecture, platform, software, and product teams to characterise performance, deploy system configurations, and improve operational processes.
Experience Level
Senior-level engineering/operations role. Prior demonstrable experience managing HPC systems is required.
Responsibilities
Primary on-site and remote responsibilities to keep HPC clusters functional, performant, and reproducible.
- Ensure uptime and baseline performance of HPC clusters in assigned regions (Germany and USA) and support customers/partners.
- Install, configure, and maintain EPYC CPU and Instinct GPU systems, BIOS and Linux tuning.
- Characterise application performance on CPU and GPU platforms and document results.
- Test early CPU samples and provide field feedback to engineering and product teams.
- Develop and drive automation to reduce repetitive operational tasks.
- Create training materials and deliver training to customers and internal teams; act as a specialist and mentor.
- Respond to and support Field Application Engineers, customers, and OEM partners to demonstrate optimal system configurations.
- Travel domestically and internationally as required (approximately 10%).
Requirements
Must-have technical skills and experience; additional items listed as nice-to-have.
-
Must-have: Clear, demonstrable track record managing production HPC systems and middleware.
- Strong Linux administration experience (Red Hat, Ubuntu, SLES).
- Knowledge of HPC ecosystem: toolchains, parallel file systems, high-performance networks, and batch managers.
- Ability to install and maintain compilers and math libraries (C/C++/Fortran, gcc, Intel OneAPI, etc.).
- Familiarity with networking technologies such as InfiniBand, Slingshot, Omnipath/BXI.
- Experience with parallel filesystems (Lustre, BeeGFS, GPFS, WekaFS) and batch managers (SLURM, PBS) and cluster management tools (Bright/TrinityX).
- Strong verbal and written communication skills in business-level English; customer-facing experience and technical documentation ability.
- Practical troubleshooting mindset and ability to work across international teams.
-
Nice-to-have: Experience building HPC applications on CPUs and GPUs, Red Hat Certified Engineer, container/orchestration knowledge (OpenStack, OpenShift, Kubernetes, Docker), CUDA familiarity, assembly-level understanding, memory/cache hierarchy profiling, HPC dataflow expertise, or security clearance.
Education Requirements
Bachelor's degree in a technical field preferred (Computer Science, Electrical Engineering, Physics, Mathematics). Equivalent practical experience is acceptable if it demonstrates the required technical skills and HPC systems expertise.
About the Company
Company: Advanced Micro Devices
Headquarters: Sunnyvale, California, USA
Advanced Micro Devices, or AMD, is a global semiconductor company that designs and manufactures microprocessors, graphics processors, and related technologies for a variety of computing devices. Known for pushing the boundaries of innovation, AMD's mission is to deliver high-performance computing solutions for AI, data centers, gaming, and embedded applications. They foster a collaborative, inclusive culture focused on creativity and problem-solving, aiming to drive progress and excellence in technology.

Date Posted: 2026-08-19