FewerJobs.
All jobs

Senior Machine Learning Engineer - Model Evaluations, Public Sector

Scale AI - San Francisco, CA; St. Louis, MO; New York, NY; Washington, DC

Posted Jun 7, 2026

Benefits

Parental leave
Not verified
Non-birth-parent leave
Not verified
Family-building benefits
  • Fertility benefits: Not verified
  • Adoption assistance: Not verified
  • Surrogacy assistance: Not verified
Mental health support
Not verified
Relocation assistance
Not verified
Childcare support
Not verified
Learning budget
Not verified
Verification
Not verified
Salary
$216K-$300K From the posting source
401(k) match
Not verified

Was this benefit information wrong? Tell us.

Market context

U.S. role benchmark (BLS OEWS)
$111,944 U.S. median for this role
Projected growth (BLS Employment Projections)
+13.7% - Much faster than average

131% above the BLS role benchmark for data and ml aggregate.

Matched to SOC 15-1252 - Data and ML aggregate by role bucket.

Source: U.S. Bureau of Labor Statistics, OEWS, May 2024 and Employment Projections, 2024-2034.

Role

Role function
Data From the posting source
Seniority
Senior From the posting source

Schedule

Shift type
Not verified
Weekend work
Not verified

Company

Equity
Offered From the posting source

Application

Cover letter
Not verified
Assessment
Not verified
Deadline
Not stated

Where they hire

State eligibility is not yet verified.

About this role

Senior Machine Learning Engineer - Model Evaluations, Public Sector San Francisco, CA; St. Louis, MO; New York, NY; Washington, DC Senior Machine Learning Engineer - Model Evaluations, Public Sector The Public Sector ML team at Scale deploys advanced AI systems-including LLMs, agentic models, and multimodal pipelines-into mission-critical government environments. We build evaluation frameworks that ensure these models operate reliably, safely, and effectively under real-world constraints. As an ML Engineer, you will design, implement, and scale automated evaluation pipelines that help customers trust and operationalize advanced AI systems across defense, intelligence, and federal missions. You will: Develop and maintain automated evaluation pipelines for ML models across functional, performance, robustness, and safety metrics, including LLM-judge-based evaluations. Design test datasets and benchmarks to measure generalization, bias, explainability, and failure modes. Build evaluation frameworks for LLM agents, including infrastructure for scenario-based and environment-based testing. Conduct comparative analyses of model architectures, training procedures, and evaluation outcomes. Implement tools for continuous monitoring, regression testing, and quality assurance for ML systems. Design and execute stress tests and red-teaming workflows to uncover vulnerabilities and edge cases. Collaborate with operations teams and subject matter experts to produce high-quality evaluation datasets. Comfortable with light travel (approximately 10%) for customer interaction and team needs. This role will require an active security clearance or the ability to obtain a security clearance. Ideally you'd have: Experience in computer vision, deep learning, reinforcement learning, or NLP in production settings. Strong programming skills in Python; experience with TensorFlow or PyTorch. Background in algorithms, data structures,

Read the full description at job-boards.greenhouse.io. FewerJobs shows a preview and links to the original posting.

Apply at job-boards.greenhouse.io

Apply link not verified; last-live date unavailable.

What verified means

Verified means a displayed claim has field-level provenance to a source FewerJobs pulled: a government or employer source, or the original job posting. Posting-sourced facts are employer-stated and are labeled separately from government records.

Related jobs