Job Description
The Company
Toters is an on-demand e-commerce and delivery platform that enables customers to get anything in their city with the highest level of convenience.
Technology is at the heart of everything we do. Our product and engineering teams work every day to build experiences that make customers’ lives easier while continuously improving internal systems to deliver faster and at the best cost.
If you’re excited about working in a high-growth startup environment and want to be part of a team shaping the future of how people shop in the Middle East, we’d love to hear from you.
About the Role
AI agents are changing both what we build and how we build it. At Toters, we are building an AI platform that supports both product-facing agents and the engineers who use AI coding agents to build software.
As one of the first engineers on this platform, you will help build the foundations that make AI agents safe, measurable, scalable, and cost-effective. You’ll work across both product agents and AI-assisted software development, go deep in at least one area, and help shape what we build versus buy.
This is a platform engineering role. It is not a prompt-writing role, and it is not a research role.
In this role, you will:
Build shared AI platform foundations
- Build and operate the LLM gateway, including routing, fallback, PII redaction, guardrails, and per-team cost attribution.
- Build and operate the MCP gateway, including tool registration, authentication, and making internal APIs agent-ready through configuration rather than bespoke code.
- Build a unified evaluation backbone for production and coding agents, including trajectory evaluations, tool-call correctness, LLM-as-judge calibration, regression suites, and OpenTelemetry tracing.
- Provide isolated, fast-start execution environments where agents can safely perform real work.
- Build and maintain a curated registry of skills, tools, and instructions with linting, versioning, and usage telemetry.
- Track AI cost per task and per merged PR, reduce context bloat, and make AI spend visible by team.
Drive AI-assisted software development
- Own the platform and harness around coding agents such as Claude Code, Cursor, and other emerging tools.
- Define instruction layering, skill distribution, and plan/execute strategies for coding agents.
- Move checks into the development inner loop and introduce agent-aware CI stages, including evaluations, AI review, and self-healing for routine failures.
- Run background agents for migrations, upgrades, test coverage, and technical cleanup.
- Define and enforce an LLM-ready repository standard across our Go, Java, and PHP codebases.
- Measure AI adoption and its impact on cycle time, change failure rate, DORA metrics, and cost.
- Drive enablement through pairing, champions, and workflow improvements rather than training alone.
Build product agent capabilities
- Build a framework-agnostic agent SDK covering planning loops, tool use, state, and memory.
- Standardize observability and deployment for AI agents.
- Implement agent identity and access controls so every agent action can be traced back to a human, with a complete audit trail.
- Partner with Product, ML, and Backend teams to take AI agent use cases from prototype to production on a standardized platform.
Your First Six Months
- Establish baseline AI usage, engineering cycle time, and AI spend.
- Launch the first version of the coding-agent harness and LLM-ready repository standard on pilot repositories.
- Stand up the LLM gateway.
- Introduce agent-aware CI stages and the first version of the evaluation backbone.
- Launch the first background-agent workflow through normal CI and human review.
- Expand the repository standard across the engineering organization.
- Support the launch of the first production product agent on the standardized platform, with release-gating evaluations and cost attribution.
Key Qualifications
Senior
- 6+ years of experience in backend, platform, DevEx, or infrastructure engineering, working on systems at meaningful scale.
- 1+ year of experience shipping LLM-based systems or agentic tooling to production.
- Deep expertise in at least one of the AI Platform tracks, with the ability to contribute to shared platform foundations.
Staff
- 9+ years of experience, including ownership of platforms or tooling used by other engineers on a daily basis.
- 2+ years running LLM systems in production, including building abstractions over model providers.
- Strong experience setting technical direction and influencing how teams work without direct management responsibility.
- Deep expertise across the shared AI platform foundations and at least one of the product-agent or AI SDLC tracks.
Both levels
- Strong experience with Go or Java and comfortable using Python for AI tooling.
- Strong distributed-systems knowledge, including gateways and proxies, authentication, rate limiting, and failure handling in latency-sensitive systems.
- Hands-on experience designing and implementing agent evaluations beyond manual spot-checking.
- Daily, hands-on experience with agentic coding tools and a strong understanding of context engineering and common agent failure modes.
- Strong CI/CD knowledge, including pipeline design, testing strategies, and build performance.
- Experience with AWS, Kubernetes, and Terraform or equivalent technologies.
- Strong technical judgment and the ability to distinguish production-ready solutions from concepts that only work in demos.
Nice to Have
- Experience building or operating MCP servers or gateways; familiarity with A2A.
- Experience with background agents or large-scale automated migration systems.
- Experience with AI evaluation and observability tools such as Langfuse, LangSmith, Braintrust, or Weave.
- Experience with developer productivity measurement frameworks such as DORA, SPACE, or Core 4.
- Experience with agent identity, access control, or Zero Trust patterns.
- Experience with RAG, search, or knowledge graphs for grounding agents in company data.
- Cross-functional experience across platform engineering, DevEx, solutions engineering, or forward-deployed engineering.
What Success Looks Like
- Approximately 100% of model calls go through the governed AI gateway.
- Production agents have release-gating evaluations.
- 60%+ of active repositories meet the LLM-ready standard.
- A clear majority of engineers use agentic tools regularly.
- Agent-authored PRs increase while change failure rates remain flat or improve.
- Engineering cycle time measurably improves against the pre-rollout baseline.
- AI spend is visible by team, task, and merged PR.
- Teams can move from an AI agent idea to a governed proof of concept in days rather than weeks.
Why Toters?
- Flexible work environment with hybrid-friendly roles.
- Opportunity to build a new AI platform from the ground up and shape our engineering strategy.
- Work across Product, ML, Backend, Infrastructure, and Engineering teams.
- Opportunity to influence how engineers across Toters build software with AI.
- Strong culture of mentorship, collaboration, and continuous learning.
- Direct impact on products and systems used by thousands of customers every day.
- Competitive compensation package.
- Exclusive discounts on Toters orders.
- First-class medical insurance.