Software Development Engineer, EC2 UltraServer Availability
Amazon - Seattle, Washington, USA
Posted May 13, 2026
Benefits
- Parental leave
- Not verified not verified - source not recorded; timestamp not recorded
- Non-birth-parent leave
- Not verified not verified - source not recorded; timestamp not recorded
- Family-building benefits
-
- Fertility benefits: Not verified
- Adoption assistance: Not verified
- Surrogacy assistance: Not verified
- Mental health support
- Not verified
- Relocation assistance
- Not verified
- Childcare support
- Not verified
- Learning budget
- Not verified
- Verification
- Not verified
- Salary
- Not verified not verified - source not recorded; timestamp not recorded
- 401(k) match
- Not verified
Was this benefit information wrong? Tell us.
Schedule
- Shift type
- Not verified
- Weekend work
- Not verified
Application
- Cover letter
- Not verified
- Assessment
- Not verified
- Deadline
- Not stated
Where they hire
State eligibility is not yet verified.
About this role
Software Development Engineer, EC2 UltraServer Availability Seattle, Washington, USA The Software Development Engineer II will design, build, and maintain cloud-based repair and recovery workflows for NVIDIA GB200 / GB300 UltraServers, orchestrating repair and recovery operations from impairment detection through completed recovery. This role requires expertise in AWS services, system architecture, and cross-functional collaboration with Capacity Management, Hardware Engineering, and Datacenter Operations to manage AI/ML infrastructure. Key job responsibilities The Software Development Engineer (SDE II) on the EC2 UltraServer Availability team is responsible for ensuring high availability of customer GB200 and GB300 UltraServers by orchestrating complex repair and recovery workflows. Following are the core responsibilities System Design & Architecture * Design and architect solutions that are cross-functional to Capacity Management, Hardware Engineering, and Datacenter Operations * Work in environments where the technology strategy is defined but the solution design is not * Build solutions that are stable, logical, testable, and efficient with the ability to independently make trade-off decisions * Investigate and develop design concepts to frame solution sets at an application and product level Software Development * Build cloud-based solutions using AWS native services for scaling infrastructure frameworks * Write high-quality, maintainable code with proper testing and code reviews * Develop and maintain the repair and recovery workflows for GB200 and GB300 UltraServer hosts * Implement automation for diagnostic triage, hardware testing, cable validation, and testing processes * Create observable systems with appropriate metrics and alarming Operational Excellence * Execute and monitor UltraServer workflows for UltraServer repair * Troubleshoot workflow
Read the full description at www.amazon.jobs. FewerJobs shows a source-linked preview and links to the original posting.
Apply link not verified; last-live date unavailable.
What verified means
Verified means a displayed claim has a recorded source field, a source URL when available, and a timestamp showing when FewerJobs checked or enriched the evidence.
Related jobs
-
Systems Engineer - (Execution) - Level 3/4
Northrop Grumman - United States-Alabama-Huntsville
-
Business Analyst (Top Secret cleared)
ICF International INC - Washington, DC
-
Engineering Project Specialist II (Full Time) - United State
Cisco - San Jose, California, US
-
Automation AI Ops Engineer
Cisco - 2 Locations