Software Development Engineer AI/ML, Inference Serving, AWS Neuron
Amazon - Cupertino, California, USA
Posted Sep 19, 2025
Benefits
- Parental leave
- 6 weeks From the posting source checked Jun 20, 2026
- Non-birth-parent leave
- 6 weeks From the posting source checked Jun 20, 2026
- Family-building benefits
- Mental health support
- Offered From the posting source checked Jun 20, 2026
- Relocation assistance
- Not verified
- Childcare support
- Offered From the posting source checked Jun 20, 2026
- Learning budget
- Not verified
- Verification
- Source-linked checked Jun 7, 2026
- Salary
- $193K-$262K From the posting source checked Jun 20, 2026
- 401(k) match
- Reported from DOL Form 5500 industry filing (not employer-specific)
Was this benefit information wrong? Tell us.
Market context
- U.S. role benchmark (BLS OEWS)
- $116,543 U.S. median for this role
- Projected growth (BLS Employment Projections)
- +9.8% - Much faster than average
95% above the BLS role benchmark for software engineering aggregate.
Matched to SOC 15-1252 - Software Engineering aggregate by role bucket.
Source: U.S. Bureau of Labor Statistics, OEWS, May 2024 and Employment Projections, 2024-2034.
Role
Schedule
- Shift type
- Not verified
- Weekend work
- Not verified
Company
Application
- Cover letter
- Not verified
- Assessment
- Not verified
- Deadline
- Not stated
Where they hire
State eligibility is not yet verified.
About this role
Software Development Engineer AI/ML, Inference Serving, AWS Neuron Cupertino, California, USA AWS Neuron is the software stack powering AWS Inferentia and Trainium machine learning accelerators, designed to deliver high-performance, low-cost inference at scale. The Neuron Serving team develops infrastructure to serve modern machine learning models-including large language models (LLMs) and multimodal workloads-reliably and efficiently on AWS silicon. We are seeking a Software Development Engineer to lead and architect our next-generation model serving infrastructure, with a particular focus on large-scale generative AI applications. Key job responsibilities * Architect and lead the design of distributed ML serving systems optimized for generative AI workloads * Drive technical excellence in performance optimization and system reliability across the Neuron ecosystem * Design and implement scalable solutions for both offline and online inference workloads * Lead integration efforts with frameworks such as vLLM, SGLang, Torch XLA, TensorRT, and Triton * Develop and optimize system components for tensor/data parallelism and disaggregated serving * Implement and optimize custom PyTorch operators and NKI kernels * Mentor team members and provide technical leadership across multiple work streams * Drive architectural decisions that impact the entire Neuron serving stack * Collaborate with customers, product owners, and engineering teams to define technical strategy * Author technical documentation, design proposals, and architectural guidelines A day in the life You'll lead critical technical initiatives while mentoring team members. You'll collaborate with cross-functional teams of applied scientists, system engineers, and product managers to architect and deliver state-of-the-art inference capabilities. Your day might involve: * Leading
Read the full description at www.amazon.jobs. FewerJobs shows a preview and links to the original posting.
Apply link not verified; last-live date unavailable.
What verified means
Verified means a displayed claim has field-level provenance to a source FewerJobs pulled: a government or employer source, or the original job posting. Posting-sourced facts are employer-stated and are labeled separately from government records.
Related jobs
-
Staff Software Engineer - ML Observability
Datadog - Boston, Massachusetts, USA; New York, New York, USA
-
Technical Leader, Software Engineer
Cisco - Milpitas, California, US
-
Senior AI Operations (AI Ops) Engineer
Navan INC - Palo Alto, CA
-
Sr. Software Engineer I, Applied AI
Axon Enterprise - Seattle, Washington, United States