Job Description
Job Summary
Description
In this role, you will:
- Build and operate AI evaluation workflows that measure the quality of LLM outputs across chat,
summarization, recommendations, and agent-based features.
- Implement LLM-as-a-Judge and rubric-based evals to score outputs for correctness, relevance, grounding,
and consistency.
- Instrument LLM and agent workflows to capture traces, prompts and responses, metadata, and user
feedback.
- Support release readiness by running evals before launches and highlighting regressions or quality risks.
- Help define agent-specific evals (task completion, tool correctness, error recovery).
- Partner with AI engineers and AI platform teams to translate product requirements into evals criteria.
- Review eval results and recommend improvements.
- Contribute to system design for observability, retries, and logging.
Key Responsibilities
Minimum Qualifications
- 4+ years of experience in data and AI-related fields such AI engineering, software development, ML
engineering, data science, or QA roles.
- We’re looking for someone with an eagerness and ability to learn new skills and solve dynamic problems
in an encouraging and expansive environment.
- Working across global teams to ensure alignment of product development.
- Strong Python skills.
- Applied knowledge of GenAI and RAG strategies, micro-services, recommendation systems, and context
engineering.
- Familiarity of AI evaluation techniques, such as Golden datasets, LLM-as-judge, or rubric-based scoring.
- Experience with different LLM ecosystems (OpenAI, Anthropic, Gemini, etc.), RAG pipelines, vector
databases (e.g., Pinecone, FAISS, Milvus, PostgreSQL).
- Proficiency in SQL and experience with at least one major data analytics platform, such as Hadoop,
Spark, or Snowflake.
- Experience with CI/CD or release validation workflows.
- Familiarity with telemetry and evaluation frameworks for AI agents.
- Experience working with data science teams on insights generation leveraging LLMs.
- Knowledge of project management, and productivity tools such as Wrike and Miro.
- Strong time management skills with the ability to collaborate across multiple teams.
- Able to balance competing priorities, long-term projects, and ad hoc requirements.
- Ability to work in a fast-paced, dynamic, constantly evolving business environment.
Preferred Qualifications
- Hands-on experience with Langfuse or similar tools for LLMs observability.
- Sound communication skills - expert at messaging domain and technical content, at a level appropriate
for the audience. Strong ability to gain trust with stakeholders and senior leadership.
- Familiarity with embeddings, retrieval algorithms, agents, and data modeling for vector development
graphs.
- Other complementary technologies for distributed systems architecture and asynchronous messaging,
agent communication, and catching like RabbitMQ, Redis, and Valkey are preferred.
Skill Requirements
Minimum Qualifications
- 4+ years of experience in data and AI-related fields such AI engineering, software development, ML
engineering, data science, or QA roles.
- We’re looking for someone with an eagerness and ability to learn new skills and solve dynamic problems
in an encouraging and expansive environment.
- Working across global teams to ensure alignment of product development.
- Strong Python skills.
- Applied knowledge of GenAI and RAG strategies, micro-services, recommendation systems, and context
engineering.
- Familiarity of AI evaluation techniques, such as Golden datasets, LLM-as-judge, or rubric-based scoring.
- Experience with different LLM ecosystems (OpenAI, Anthropic, Gemini, etc.), RAG pipelines, vector
databases (e.g., Pinecone, FAISS, Milvus, PostgreSQL).
- Proficiency in SQL and experience with at least one major data analytics platform, such as Hadoop,
Spark, or Snowflake.
- Experience with CI/CD or release validation workflows.
- Familiarity with telemetry and evaluation frameworks for AI agents.
- Experience working with data science teams on insights generation leveraging LLMs.
- Knowledge of project management, and productivity tools such as Wrike and Miro.
- Strong time management skills with the ability to collaborate across multiple teams.
- Able to balance competing priorities, long-term projects, and ad hoc requirements.
- Ability to work in a fast-paced, dynamic, constantly evolving business environment.
Preferred Qualifications
- Hands-on experience with Langfuse or similar tools for LLMs observability.
- Sound communication skills - expert at messaging domain and technical content, at a level appropriate
for the audience. Strong ability to gain trust with stakeholders and senior leadership.
- Familiarity with embeddings, retrieval algorithms, agents, and data modeling for vector development
graphs.
- Other complementary technologies for distributed systems architecture and asynchronous messaging,
agent communication, and catching like RabbitMQ, Redis, and Valkey are preferred.
Other Requirements