
Tessera Labs is a new category of enterprise software: an AI platform that changes how the world's largest companies run.
Every large enterprise carries the same weight — decades of accumulated process, data, and code that no longer match the business it has become. Changing any of it is a program measured in years and hundreds of millions of dollars, staffed by armies of consultants, and it fails more often than anyone admits. Most companies have quietly accepted this as the cost of being large.
We don't. Tessera is a transformation engine: a governed, multi-agent platform that understands an enterprise's process, data, and code as one connected system and changes it in weeks rather than years. We're vendor-agnostic by design — SAP, Salesforce, Workday, Oracle, Snowflake, MuleSoft — and tied to none of them.
Two things make this hard, and they're the reason the research is interesting. Governance: every action is logged, traceable, and reversible, because our customers are regulated and these are the systems that close their books. And generality: the platform has to work on landscapes it has never seen, at companies whose complexity is genuinely unique to them.
We sell a product, not a service. Our people are here to make the product successful, not the other way around — which is also why research here is a durable investment rather than a line item on an engagement.
We raised a $60M Series A led by Andreessen Horowitz, with Foundation Capital, Myriad Venture Partners, and Osage University Partners participating.
Building this platform has two halves — the systems that surround the model, and the model itself. This role owns the second, at scale.
Frontier models have never seen most of what we work on. Proprietary dialects, customer-specific configuration, semantics that exist in no public corpus, and tasks that run forty steps before anything tells you whether you were right. Closing that gap with post-training, environments, and evaluation is a research problem, and Research Engineers are the people who make it happen at scale.
You build the training, environment, evaluation, and inference machinery that turns a hypothesis about agent behavior into a measured result, and then into a model that ships. You'll work closely with Research Scientists and own experiments of your own within weeks. The division isn't seniority or idea ownership — Research Engineering owns the machinery and is accountable for it working at scale; Research Science owns the agenda and is accountable for the result being true.
One property makes this an unusually good RL setting: much of our task space is verifiable A transformation either produces a system that builds, passes the customer's regression suite, and behaves equivalently, or it doesn't. That's a real reward signal rather than a preference model, and building the environments that make it cheap and trustworthy is a large part of this job.
We post-train open-weight models on rented clusters. We're compute-constrained relative to a frontier lab and we buy more when a result justifies it — worth knowing up front.
Build and scale the post-training stack: SFT, preference optimization, and reinforcement learning for long-horizon tool use, transformation, and reconciliation over enterprise systems. RL is the center of gravity of this role, not a side interest.
Build the memory and context machinery that long-horizon agents run on: what an agent retains across a forty-step run, how it's structured, retrieved, compacted, and revised — and how you train a model to use it well rather than bolting it on at inference.
Build the representation layer that agents reason over — ontologies and knowledge graphs derived from real enterprise systems — and the pipelines that construct, validate, and keep them current.
Design and implement data generation and curation pipelines — synthetic landscapes, transformation traces, tool-call trajectories, curriculum infrastructure — that teach models to operate systems no public model has seen.
Build the RL environments: sandboxed landscapes and execution-and-verification harnesses where a change can be applied, checked, and scored automatically, backed by design-partner traces where synthetic data won't do.
Build and run the offline eval harness for long-horizon agentic behavior — trajectory-level scoring, task suites, and the infrastructure that makes a result reproducible six weeks later. You own the harness; the methodology it implements is a shared argument with Research Science.
Run experiments end to end — design, launch, debug, analyze — and be honest about which effects are real and which are noise.
Optimize training and inference throughput: kernels, parallelism strategies, memory, batching, serving. Long context matters here more than most places; a single enterprise artifact can eat a context window.
Take a training result from "the eval moved" to "it's serving traffic" — quantization, serving configuration, rollback path.
Establish standards for reproducibility, experiment tracking, and result hygiene, so findings survive contact with the next person who builds on them.
Building a synthetic landscape generator that produces enterprise systems with realistic complexity and coupling, then measuring what training on it actually buys on held-out real customer environments.
Standing up a distributed RL loop where reward comes from build-and-regression outcomes, and catching that the environment was leaking target state into the observation — the policy was scoring 0.9 by reading the answer rather than doing the task.
Post-training a mid-size open-weight model to match frontier-model accuracy on our core transformation tasks at a fraction of the serving cost.
Halving end-to-end latency on a multi-agent run: prefix KV-cache reuse across sub-agents so we stop re-prefilling the same schemas forty times, plus speculative decoding to shorten each decode step.
Rebuilding the eval harness so a six-month-old result can be reproduced from a commit hash.
Have significant experience training, fine-tuning, or post-training language models, and can point to results you owned. RL tuning experience — RLHF, RLAIF, RLVR, GRPO-family, or agentic RL — is close to a requirement rather than a bonus.
Have thought hard about memory and context for long-running agents, whether through architecture, retrieval, or training.
Have strong software engineering fundamentals — the code you write for experiments is code other people can run.
Are fluent in Python and PyTorch (or JAX), and comfortable debugging distributed training when the loss curve does something inexplicable.
Can design, run, and interpret an experiment with real empirical rigor. You look at data and distinguish an effect from noise from a bug.
Have worked with GPU infrastructure at scale and know where the time and the memory go.
Want your research to end up in production, and treat that as the interesting constraint rather than the tax.
Communicate clearly in writing. We make decisions from documents.
Experience building RL environments, execution sandboxes, or verifiable-reward task suites.
Experience with long-context modeling — extension, efficient attention, position methods, or the evaluation that shows whether a long-context claim is real.
Experience with knowledge graphs, ontologies, or semantic layers over structured enterprise data.
Experience with agent memory systems, episodic or otherwise, in production rather than in a paper.
Experience with code models: repository-scale context, program synthesis, automated repair, or transpilation.
Contributions to open-source ML systems (vLLM, SGLang, PyTorch, Triton, DeepSpeed, Ray, Megatron, TRL, or similar).
A track record of publications, technical reports, or open-source releases — we treat these as equivalent currency.
An advanced degree in CS, ML, math, physics, or a related quantitative field, or equivalent industry research experience.
We're looking for someone who enjoys building reliable, scalable systems that help teams move faster. If you take pride in cutting-edge AI, value clear ownership, and want to have a meaningful impact in a fast-moving environment, you’ll fit right in.
No third party may recruit, solicit candidates, publish job opportunities, use Tessera Labs’ name or branding, or represent that they are acting on behalf of Tessera Labs without prior written authorization. Any such activity conducted without explicit written consent is strictly prohibited.

Enterprise transformations shouldn't take years or cost fortunes. At Tessera, we've built a multi-agent AI platform that cuts ERP transformation timelines from years to weeks and reduces costs by more than half—while delivering first-time-right outcomes with enterprise-grade security and governance.
We combine the dependability of a trusted integrator with the speed of an AI innovator. Our vendor-agnostic platform is pre-trained on thousands of enterprise landscapes and hundreds of years of expertise, so it adapts to your environment from day one—harmonizing systems and data to deliver secure, governed execution across ERP and other critical systems.
Unlike traditional system integrators, ERP vendor tools, AI point solutions, or in-house builds, Tessera delivers transformations that compress timelines by 90%, replace people-heavy manual work with intelligent automation, ensure systems and data harmonize correctly from the start, and avoid vendor lock-in through adaptive intelligence that evolves with your workflows.
Tessera creates lasting change in how your processes and systems operate together. As the platform adapts to your environment and evolves with new data, you gain the confidence to modernize with less risk—freeing up resources to reinvest in growth rather than maintaining legacy complexity.
Proven in regulated industries. Built for enterprise reality. Reshaping how organizations transform.