Software Development Manager, LLM Inference Model Enablement, Neuron SDK
Amazon - Cupertino, California, USA
Posted Sep 5, 2025
Benefits
- Parental leave
- 6 weeks From the posting source checked Jun 20, 2026
- Non-birth-parent leave
- 6 weeks From the posting source checked Jun 20, 2026
- Family-building benefits
- Mental health support
- Offered From the posting source checked Jun 20, 2026
- Relocation assistance
- Not verified
- Childcare support
- Offered From the posting source checked Jun 20, 2026
- Learning budget
- Not verified
- Verification
- Source-linked checked Jun 7, 2026
- Salary
- $213K-$288K From the posting source checked Jun 20, 2026
- 401(k) match
- Reported from DOL Form 5500 industry filing (not employer-specific)
Was this benefit information wrong? Tell us.
Market context
- U.S. role benchmark (BLS OEWS)
- $116,543 U.S. median for this role
- Projected growth (BLS Employment Projections)
- +9.8% - Much faster than average
115% above the BLS role benchmark for software engineering aggregate.
Matched to SOC 15-1252 - Software Engineering aggregate by role bucket.
Source: U.S. Bureau of Labor Statistics, OEWS, May 2024 and Employment Projections, 2024-2034.
Role
Schedule
- Shift type
- Not verified
- Weekend work
- Not verified
Company
- Equity
- Offered Verified - SEC 10-K source checked Jun 20, 2026
Application
- Cover letter
- Not verified
- Assessment
- Not verified
- Deadline
- Not stated
Where they hire
State eligibility is not yet verified.
About this role
Software Development Manager, LLM Inference Model Enablement, Neuron SDK Cupertino, California, USA DESCRIPTION AWS Utility Computing (UC) provides product innovations, from foundational services such as Amazon Elastic Compute Cloud (EC2), to new product innovations that continue to set AWS's services and features apart in the industry. We develop AWS Neuron, the complete software stack for Trainium, Amazon's custom cloud-scale machine learning accelerators. Come optimize LLMs such as Llama and GPT-OSS to run really fast on Trainium. As the SDM for the LLM Inference Model Enablement team, you will lead a team of expert AI/ML engineers to onboard and optimize state-of-the-art open-source and customer LLMs, both dense and MoE, for inference on Neuron and Trainium and Inferentia accelerators. You will also drive improvements in model enablement speed and experience, while advancing inference usability and quality through inference features, infrastructure optimization, tools, and automation. The ideal candidate will have a strong background in LLM model architectures, model performance optimizations, and inference techniques, such as delivering high-performance models using distributed inference libraries. You should be capable of managing demanding, fast-changing priorities. You should have a strong technical ability to understand and deliver as part of a vertically integrated system stack consisting of the PyTorch inference library, Neuron compiler, runtime, and collectives. A day in the life You will work with your senior management and technical leaders to define the model enablement and performance optimization for the latest SOTA LLMs, build and deliver them to customers. Meanwhile, lead the team to continue improving the
Read the full description at www.amazon.jobs. FewerJobs shows a preview and links to the original posting.
Apply link not verified; last-live date unavailable.
What verified means
Verified means a displayed claim has field-level provenance to a source FewerJobs pulled: a government or employer source, or the original job posting. Posting-sourced facts are employer-stated and are labeled separately from government records.
Related jobs
-
Senior AI/ML Platform Engineer (LLM/SLM Inference)
Cisco - San Jose, California, US
-
Staff Software Engineer - ML Observability
Datadog - Boston, Massachusetts, USA; New York, New York, USA
-
Forward-Deployed Test Engineer
Ranger Energy Services INC - San Francisco, CA
-
Manager, Software Engineering
Union Bankshares INC - Bellevue, WA Remote