At Iris.ai, we’re building an agentic AI platform that scales expert-level domain knowledge across entire organizations.
For more than a decade, we’ve worked at the intersection of scientific research, industrial data, and applied AI, helping researchers, engineers, and business teams reason over complex technical knowledge.
Our products - Neuralith, Axion, and RSpace - span the full GenAI lifecycle:
What makes us different: we care deeply about accuracy, evaluation, and responsibility. We don’t optimize for demos and proof-of-concepts we optimize for systems that experts trust and use.
We’re looking for a Senior/Lead LLM Researcher to lead Iris.ai’s research around LLM evaluation, model interpretability, and mechanistic interpretation of LLMs.
You’ll define and drive research into how we evaluate, understand, and measure the reliability of LLMs - from answer quality and groundedness to uncertainty, confidence, and model behaviour.
This is a research and leadership role with a strong applied focus. You’ll shape research directions, lead experimentation, and work closely with our engineering and product teams to turn research into products used within the Iris.ai platform.
You’ll also participate in shaping projects for EU and national research funding, leading and co-authoring grant proposals such as Horizon Europe and EIC.
You’ll work on a focused set of high‑impact research directions that sit at the core of modern applied NLPMLLLM and agentic systems. You will be making LLM-based systems measurable, interpretable, and trustworthy. The core aspects include:
Your goal will be turning rigorous research into capabilities that real users can trust and use.
If you want to do meaningful NLP work, help secure funding for frontier AI research, and grow in a culture built on trust, rigor, and fairness — let’s talk.
We’re not your typical tech company. We believe in:
(Just imagine: Someone once bought a Tesla option for $1 — it's worth $400 today.)
We’ve built our benefits to reflect how we work: with trust, fairness, and room to grow.
If you care about building high-quality, ethical AI — guided by data and human judgment — you’ll feel at home at Iris.ai.
👉 Apply now or reach out with questions. We’re transparent by default.

Iris.ai is the AI Development and Operation Platform for building secure, high-performance Agentic RAG systems.
Built for innovation teams, AI platform leads, and R&D departments, Iris.ai helps organizations move beyond prototypes and into production with measurable results.
Our modular tools, including Neuralith, Axion, and RSpace™, transform unstructured, siloed data into agent-ready knowledge. Enterprises use Iris.ai to connect internal and external data, orchestrate domain-specific agents, and evaluate LLMs with 30+ performance, safety, and cost metrics.
Deployment is secure and flexible: on-premise, cloud, or hybrid. Governance is built in — with full data separation, privacy-by-design architecture, and ISO27001-certified infrastructure.
Trusted by organizations like ArcelorMittal, L’Oréal, USDA and the Finnish Food Authority, Iris.ai has processed over 160M documents and delivered:
– 35%+ reduction in LLM usage costs
– Up to 80% acceleration in AI go-to-market
We work with AI leaders in telecom, manufacturing, public sector, and research to operationalize AI with confidence.
Backed by the European Innovation Council and grounded in a decade of deep-tech research, Iris.ai helps enterprises turn knowledge into action — securely, efficiently, and at scale.
#AgenticAI #EnterpriseAI #RAG #LLMEvaluation #AIInfrastructure