
Imagine what you could do here. At Apple, great ideas have a way of becoming extraordinary products, services, and customer experiences very quickly. Bring passion and dedication to your work, and there’s no telling what you could accomplish.
The Channel Sales AI Product Engineering team is looking for a Machine Learning Evaluation Engineer to help build and scale evaluation capabilities for our next generation of AI-powered experiences.
In this role, you will develop evaluation frameworks, datasets, tooling, and quality signals that enable teams to understand and continuously improve Generative AI and LLM-powered products. You will work closely with Machine Learning, Software Engineering, Quality Engineering, Product, Human Interface, Data Science, and domain experts to establish rigorous evaluation practices throughout the AI product lifecycle.
You will help define how we measure the quality of AI experiences across the Commerce domain, including Store AI, Shopping AI, Learning AI, Content GenAI, Conversational AI, and Platform Self-Service.
This is an opportunity to work at the intersection of machine learning, software engineering, data, and product quality, helping ensure our AI experiences are accurate, relevant, grounded, reliable, and useful for users around the world.
As a Machine Learning Evaluation Engineer, you will design and build scalable evaluation systems for LLM, Generative AI, Conversational AI, and Agentic AI products.
You will:
Experience building evaluation infrastructure for production-scale LLM or Generative AI applications.
Experience evaluating RAG, conversational systems, AI agents, personalization, recommendations, or multimodal AI.
Experience building golden datasets, regression suites, automated quality gates, and continuous evaluation pipelines.
Experience integrating ML evaluation into CI/CD and production release processes.
Experience with prompt evaluation, model comparison, experiment tracking, and AI observability.
Experience evaluating multilingual AI experiences across languages, locales, and markets.
Familiarity with responsible AI evaluation, including robustness, safety, bias, and adversarial testing.
Experience developing internal ML platforms, developer tooling, or self-service evaluation capabilities used across multiple teams.
Experience working with large-scale datasets and distributed ML or data-processing infrastructure.
Master's degree in Computer Science, Machine Learning, Artificial Intelligence, Data Science, Statistics, Electrical Engineering, or a related technical field, or equivalent industry experience.
Typically requires a minimum of 7 years of related experience in Machine Learning Engineering, ML Evaluation, Software Engineering, Data Science, Quality Engineering, or a related technical field.
Strong programming skills in Python and experience developing production-quality software, ML systems, data pipelines, or evaluation infrastructure.
Experience developing or evaluating LLMs, Generative AI, Conversational AI, NLP, recommendation systems, or other machine-learning-driven products.
Experience designing automated ML evaluation frameworks, metrics, benchmarks, datasets, or experimentation methodologies.
Understanding of modern LLM application architectures, including prompting, embeddings, retrieval-augmented generation (RAG), tool use, and agentic workflows.
Experience with model-based evaluation techniques and an understanding of the strengths and limitations of approaches such as LLM-as-a-Judge.
Experience with Human-in-the-Loop evaluation, annotation, or data-quality workflows.
Strong understanding of statistical analysis, experimentation, sampling, and measurement methodologies.
Experience performing model error analysis, failure analysis, and root-cause investigation.
Ability to work effectively across Machine Learning, Engineering, Product, Quality, and Data teams.
Excellent written and verbal communication skills, with the ability to translate complex technical findings into clear, actionable recommendations.
Bachelor's degree in Computer Science, Machine Learning, Artificial Intelligence, Data Science, Statistics, Electrical Engineering, or a related technical field, or equivalent industry experience.

We’re a diverse collective of thinkers and doers, continually reimagining what’s possible to help us all do what we love in new ways. And the same innovation that goes into our products also applies to our practices — strengthening our commitment to leave the world better than we found it. This is where your work can make a difference in people’s lives. Including your own.
Apple is an equal opportunity employer that is committed to inclusion and diversity. Visit apple.com/careers to learn more.