Software Development Manager, LLM Inference Model Enablement, Neuron SDK

Amazon - Cupertino, California, USA

Posted Sep 5, 2025

Benefits

Parental leave: Not verified not verified - source not recorded; timestamp not recorded
Non-birth-parent leave: Not verified not verified - source not recorded; timestamp not recorded
Family-building benefits: Fertility benefits: Not verified
Adoption assistance: Not verified
Surrogacy assistance: Not verified
Mental health support: Not verified
Relocation assistance: Not verified
Childcare support: Not verified
Learning budget: Not verified
Verification: Not verified
Salary: Not verified not verified - source not recorded; timestamp not recorded
401(k) match: Not verified

Was this benefit information wrong? Tell us.

Schedule

Shift type: Not verified
Weekend work: Not verified

Application

Cover letter: Not verified
Assessment: Not verified
Deadline: Not stated

Where they hire

State eligibility is not yet verified.

About this role

Software Development Manager, LLM Inference Model Enablement, Neuron SDK Cupertino, California, USA DESCRIPTION AWS Utility Computing (UC) provides product innovations, from foundational services such as Amazon Elastic Compute Cloud (EC2), to new product innovations that continue to set AWS's services and features apart in the industry. We develop AWS Neuron, the complete software stack for Trainium, Amazon's custom cloud-scale machine learning accelerators. Come optimize LLMs such as Llama and GPT-OSS to run really fast on Trainium. As the SDM for the LLM Inference Model Enablement team, you will lead a team of expert AI/ML engineers to onboard and optimize state-of-the-art open-source and customer LLMs, both dense and MoE, for inference on Neuron and Trainium and Inferentia accelerators. You will also drive improvements in model enablement speed and experience, while advancing inference usability and quality through inference features, infrastructure optimization, tools, and automation. The ideal candidate will have a strong background in LLM model architectures, model performance optimizations, and inference techniques, such as delivering high-performance models using distributed inference libraries. You should be capable of managing demanding, fast-changing priorities. You should have a strong technical ability to understand and deliver as part of a vertically integrated system stack consisting of the PyTorch inference library, Neuron compiler, runtime, and collectives. A day in the life You will work with your senior management and technical leaders to define the model enablement and performance optimization for the latest SOTA LLMs, build and deliver them to customers. Meanwhile, lead the team to continue improving the

Read the full description at www.amazon.jobs. FewerJobs shows a source-linked preview and links to the original posting.

Apply at amazon.jobs

Apply link not verified; last-live date unavailable.

What verified means

Verified means a displayed claim has a recorded source field, a source URL when available, and a timestamp showing when FewerJobs checked or enriched the evidence.

Related jobs

Hardware System and Board Failure Analysis Technical Lead

Cisco - Milpitas, California, US
Sr. Staff System Architect

Northrop Grumman - United States-Illinois-Rolling Meadows
Senior Project Manager - Product Implementation

Deluxe CORP - 2 Locations
Sr. Staff Product Operations Manager, Product Lifecycle (Remote)

Cisco - Coral Gables, Florida, US

Software Development Manager, LLM Inference Model Enablement, Neuron SDK

Benefits

Schedule

Application

Where they hire

About this role

What verified means

Related jobs

Hardware System and Board Failure Analysis Technical Lead

Sr. Staff System Architect

Senior Project Manager - Product Implementation

Sr. Staff Product Operations Manager, Product Lifecycle (Remote)