Fetch

Staff AI Evaluation Lead

Fetch  •  $129k - $152k/yr  •  Remote  •  3 days ago
Apply
AI can make mistakes so check important info. Chat history is never stored.

Job Description

About the Role:
At Fetch, we’re building AI and automation systems that make our work smarter, faster, and more scalable. The AI Operations team ensures our models, automations, and LLM systems perform with quality, reliability, and measurable impact.

As a Staff AI Evaluation Lead, you’ll own automation and evaluation programs across AI Operations. You’ll translate business and functional goals into scalable systems, define how we measure quality, and ensure automation and evaluation become durable, high-impact capabilities across the organization.
This role is ideal for someone who combines deep technical problem-solving with systems-level thinking and strong cross-functional leadership.

This is a full-time role that can be held from one of our US offices or remotely in the United States.

Role Responsibilities: 
  • Own programs: Lead complex, high-impact automation and evaluation initiatives across workflows or teams
  • Design scalable solutions: Architect end-to-end workflows integrating datasets, evaluations, automations, and HITL processes; build the harness and reusable components other analysts build against; build evaluation pipelines that run against production without hands-on operation
  • Establish standards: Define dataset standards, evaluation methodology, failure taxonomy, and quality measurement across AI Operations; verify that projects you don't run are meeting them; own the process by which those standards get made and used
  • Own metric definitions: Create and maintain the metric definitions the org evaluates against, validate them against real production behavior, and revise them as models, tooling, and system architectures change
  • Set the bar for production: Define the quality bar a system clears before it reaches production and where human review stays in the loop, and pull a system back when it stops meeting the bar
  • Allocate evaluation depth by risk: Decide which systems get what depth of evaluation given finite capacity, name the risk accepted on the rest, and make that tradeoff visible to project stakeholders
  • Keep the evaluation stack current: Define when a change to a model or platform requires re-baselining across the org and own the process for doing it; evaluate new models and AI capabilities as they ship and decide what the org adopts
  • Improve systems at scale: Lead redesign of workflows, tooling, and processes to improve performance and durability across AI Operations, not only within your own programs
  • Drive cross-functional alignment: Influence priorities and partner with Engineering, Product, and AI teams to deliver solutions
  • Measure and communicate impact: Define success metrics and communicate performance and recommendations to leadership
  • Elevate the team: Raise the bar through your work and actively mentor others by sharing approaches, guiding problem-solving, and enabling the team to build stronger automation and evaluation capabilities
  • Drive innovation: Stay current on emerging tools and approaches; pilot and translate them into actionable improvements for the team

Minimum Requirements:
  • 8+ years of professional experience AI, machine learning, operational automations, or a related field.
  • Proven ability to lead complex automation or evaluation initiatives across systems or teams
  • Experience designing, building, and scaling evaluation frameworks, datasets, and quality systems at scale for LLMs or AI products
  • Fluency in SQL, JSON, APIs, and scripting with AI assistance, and ownership of the technical direction of the evaluation stack
  • Deep working knowledge of LLM and agentic system behavior, automation platforms, and system design, demonstrated in systems you have built
  • Experience with data pipelines, APIs, and production systems
  • Experience in influencing cross-functional stakeholders and aligning priorities
  • Demonstrated ability to define metrics and drive measurable business impact

Preferred Requirements:
  • Experience mentoring or leading technical contributors
  • Experience evaluating agentic systems in production at scale
  • Experience setting technical standards adopted across an organization
 
Compensation:
At Fetch, we offer competitive compensation packages including base, equity, and benefits to the exceptional folks we hire. The base salary range for this position is $129,000 - $152,000. Discover our benefits and how our employees live rewarded at https://fetch.com/careers

Fetch

About Fetch

Fetch, America's Rewards App, empowers consumers to Live Rewarded and helps brands create lifelong customers through the power of Fetch Points.

Fetch has sweeping visibility into what consumers buy, capturing more than $179 billion worth of transactions annually using cutting-edge artificial intelligence and machine learning technologies. To date, Fetch users have submitted more than 5 billion receipts and earned more than $1 billion in rewards.

The app is available to download on the App Store and Google Play Store and has more than 6 million five-star reviews from happy Fetchers.

Industry
IT & Software
Company Size
1,001-5,000 employees
Headquarters
Hybrid-remote workplace
Year Founded
2013
Website
fetch.com
Social Media