FewerJobs.
All jobs

Annotation Data Scientist, Evaluation Integrity (Siri)

Apple - Cambridge, United States of America

Posted May 19, 2026

Benefits

Parental leave
Not verified
Non-birth-parent leave
Not verified
Family-building benefits
  • Fertility benefits: Not verified
  • Adoption assistance: Not verified
  • Surrogacy assistance: Not verified
Mental health support
Not verified
Relocation assistance
Not verified
Childcare support
Not verified
Learning budget
Not verified
Verification
Not verified checked Jun 13, 2026
Salary
$155K-$275K From the posting source checked Jun 20, 2026
401(k) match
Reported from DOL Form 5500 industry filing (not employer-specific)

Was this benefit information wrong? Tell us.

Market context

U.S. role benchmark (BLS OEWS)
$111,944 U.S. median for this role
Projected growth (BLS Employment Projections)
+13.7% - Much faster than average

92% above the BLS role benchmark for data and ml aggregate.

Matched to SOC 15-1252 - Data and ML aggregate by role bucket.

Source: U.S. Bureau of Labor Statistics, OEWS, May 2024 and Employment Projections, 2024-2034.

Role

Role function
Data From the posting source checked Jun 20, 2026
Seniority
Mid From the posting source checked Jun 20, 2026

Schedule

Shift type
Not verified
Weekend work
Not verified

Company

Company stage
Public-company From the posting source checked Jun 20, 2026
Equity
Offered Verified - SEC 10-K source checked Jun 20, 2026

Application

Cover letter
Not verified
Assessment
Not verified
Deadline
Not stated

Where they hire

State eligibility is not yet verified.

About this role

Annotation Data Scientist, Evaluation Integrity (Siri) Cambridge, United States of America Play a part in the ongoing revolution in human-computer interaction. Siri is evolving - and the way we evaluate it has to evolve with it. Join the Evaluation Integrity team to help build the trusted quality signal behind every Siri release. Within the Siri evaluation organization, the Human Evaluation sub-team is responsible for answering the question: can we trust our evals? We do that by designing human-in-the-loop (HITL) annotation tasks that scrutinize every moving part of an agentic evaluation - the simulated user agent, the conversation it has with Siri, and the automated evaluators that grade the exchange. This role sits at the intersection of data science, human annotation engineering, and evaluation methodology, and is instrumental in turning human judgment into a rigorous, reproducible signal that directly informs pre-ship model and product decisions. As an Annotation Data Scientist on the Evaluation Integrity team, you will design and run HITL annotation projects that evaluate the quality and authenticity of agentic user personae, the validity of agent-to-agent conversations, and the reliability of LLM-as-judge and rule-based evaluators against Siri's product specifications. You will own annotation initiatives end-to-end; from rubric design and tooling, through annotator calibration, to data science analysis that turns annotator judgments into actionable signal for modeling, planning, and product teams. Design HITL annotation tasks for agentic evaluation. Advise on rubrics and design workflows that ask annotators to assess (a) the quality and authenticity of user agent personae, (b) the validity

Read the full description at jobs.apple.com. FewerJobs shows a preview and links to the original posting.

Apply at jobs.apple.com

Apply link not verified; last-live date unavailable.

What verified means

Verified means a displayed claim has field-level provenance to a source FewerJobs pulled: a government or employer source, or the original job posting. Posting-sourced facts are employer-stated and are labeled separately from government records.

Related jobs