
Kimchi is the AI platform inside CAST AI. We started by helping companies run LLMs on their own Kubernetes clusters and now we're providing a managed variant of those same capabilities.
Multi-model inference (MiniMax, Kimi, GLM-5, Nemotron, DeepSeek) with intelligent routing, an OpenAI-compatible API and deployment ranging from our GPUs to your own VPC. The inference layer is the foundation and the API is what sits in front of it as the primary channel for broadly and reliably distributing our AI services and powers our own Kimchi harness.
As a Senior Software Engineer, you will have the opportunity to work on different key features of our product. All of these are high-agency roles across multiple parts of the tech stack that minimize process friction that would otherwise prevent you from shipping.
In every team you will own features end-to-end: design, implementation, testing, production rollout. Most projects ship in 1-4 weeks. You'll work directly with product and other engineering teams on problems that don't have textbook solutions.
We are currently hiring Senior Software Engineers for the Harness team:
OpenAI and Anthropic ship models. They also ship one harness each – the scaffolding that turns a raw model into something that can plan, execute, recover, and complete work. We ship a different kind of harness: one built for cost-conscious, long-horizon autonomy, running on inference infrastructure we control end-to-end.
A decent model with a great harness beats a great model with a bad harness. We've watched this play out. The gap between what today's models can do and what you see them doing is largely a harness gap – and that gap is where we operate.
As part of our standard hiring process, we would like to inform you that a background check may be conducted at the final stage of recruitment through our third-party provider, Checkr.
Please note that Cast AI does not provide any form of visa sponsorship/work permit.
#LI-Remote

Increase your profit margin without additional work. CAST AI cuts your cloud bill in half, automates DevOps tasks, and prevents downtime in one Autonomous Kubernetes platform.