Sr. Technical Program Manager - AI/ML Hardware Health & Stability, Global Data Center Operations
Amazon - Seattle, Washington, USA
Posted Apr 8, 2026
Benefits
- Parental leave
- Not verified not verified - source not recorded; timestamp not recorded
- Non-birth-parent leave
- Not verified not verified - source not recorded; timestamp not recorded
- Family-building benefits
-
- Fertility benefits: Not verified
- Adoption assistance: Not verified
- Surrogacy assistance: Not verified
- Mental health support
- Not verified
- Relocation assistance
- Not verified
- Childcare support
- Not verified
- Learning budget
- Not verified
- Verification
- Not verified
- Salary
- Not verified not verified - source not recorded; timestamp not recorded
- 401(k) match
- Not verified
Was this benefit information wrong? Tell us.
Schedule
- Shift type
- Not verified
- Weekend work
- Not verified
Application
- Cover letter
- Not verified
- Assessment
- Not verified
- Deadline
- Not stated
Where they hire
State eligibility is not yet verified.
About this role
Sr. Technical Program Manager - AI/ML Hardware Health & Stability, Global Data Center Operations Seattle, Washington, USA The Central Operations team within Amazon Web Services (AWS) Infrastructure is seeking a Senior Technical Program Manager to drive the health, stability, and operational excellence of new hardware deployments across our global data center fleet. This role uniquely blends technical program management with strategic account management to ensure our GenAI and high-performance computing infrastructure delivers maximum value to customers. As a Sr. TPM, you will be the technical advocate and strategic advisor for operational support of new AI/ML hardware platforms. You will serve as the central owner of operational health (failure rate, repair efficacy, repair dwell time, break/fix process improvement) while driving cross-functional initiatives to improve these key performance indicators. You will work at the intersection of hardware engineering, data center operations, and service teams like EC2-translating complex technical data into actionable insights and leading programs that accelerate capacity delivery while maintaining the highest standards of operational health. This is not a sales role, but rather an opportunity to be the 'voice of the customer' and the 'voice of operations' for critical infrastructure that powers AWS's most demanding workloads. You will craft and execute strategies to optimize new hardware deployments, proactively identify and remediate stability issues, and establish best practices that scale across AWS's global infrastructure. Key job responsibilities Hardware Health & Stability Leadership - Own the end-to-end health and stability metrics for new AI/ML hardware platforms, establishing KPIs and routines that provide
Read the full description at www.amazon.jobs. FewerJobs shows a source-linked preview and links to the original posting.
Apply link not verified; last-live date unavailable.
What verified means
Verified means a displayed claim has a recorded source field, a source URL when available, and a timestamp showing when FewerJobs checked or enriched the evidence.
Related jobs
-
Mechanical Engineering Manager 2 - 16282
Northrop Grumman - United States-Utah-Roy
-
Senior Software Engineer, Simulation and Integration
Axcelis Technologies INC - Beverly, MA
-
Payload AI&T Lead Staff Systems Engineer
Northrop Grumman - United States-Maryland-Linthicum
-
Senior Software Engineer, Equipment Control
Axcelis Technologies INC - Beverly, MA