MoyaiOpen source cloud agent
HermesWorks with Claude Code and CodexSelf-hosted · 100+ providers through LiteLLM
- [1]The problem
- [2]The results: 79% cheaper
- [3]Why we're open sourcing it
- [4]A cloud agent that keeps working
- [5]Any harness
- [6]Any model, any provider
- [7]Get started
Today we're open sourcing Moyai, the self-hosted cloud agent our team uses every day. It works with Claude Code and Codex, runs on 100+ providers through LiteLLM, and turns a task from Slack or the browser into a pull request while your laptop is closed
The problem
Our Devin bill hit $101,872 in a single month, and it was only being used internally. Most of that spend came from sessions and automations our own engineers kicked off, on models and routing we had no control over

We already run a gateway that routes across 100+ providers. We wanted to point our coding agent at our own models and our own routing logic, and pay for inference instead of seats
The results: 79% cheaper
Moyai does the same work for about $700 a day. Over the same 31 days that's roughly $21,700 instead of $101,872, so we kept about $80,000 of a single month's bill
Why we're open sourcing it
Last week we wrote about how we built our own internal Devin in 2 days. The response was mostly one question: can we run it too?
Now you can. Moyai is the same code we run in production at LiteLLM. Deploy it on your own infrastructure, point it at your own LiteLLM gateway, and keep your code, credentials and spend inside your own accounts
A cloud agent that keeps working
Every session gets its own cloud workspace with a terminal, a filesystem and a browser. The agent edits code, runs your tests and prepares a pull request for review, all without touching anyone's laptop
Sessions are durable. You can follow along from the web app, send a correction mid-task, or pick the thread back up in Slack the next morning. Large tasks can fan out to parallel worker agents, each on its own machine, and come back together when they finish
Start a task in Slack. Come back to a PR.
Any harness
The agent loop is a choice, not a lock-in. Pick the harness for each session from the composer, and Moyai runs it in the same isolated workspace with the same tools, connections and permissions
HermesHermes is the default. Claude Code, Codex, OpenCode and Deep Agents run through the LiteLLM agent SDK, so adding the next harness is a registry entry instead of a rewrite
Any model, any provider
Every model request goes through LiteLLM. Switch from GPT-6 Astra to Claude Opus 5.5 to GLM-5.3 between messages, and every request is attributed to the teammate who made it. Provider keys stay on the server; the sandbox never sees them
That's 100+ providers out of the box. If LiteLLM can call it, Moyai can use it
Get started
Clone the repo and try the local demo in a couple of minutes, no API keys required
git clone https://github.com/BerriAI/moyai.git
cd moyai
cp .env.example .env
uv sync --frozen
uv run uvicorn app.main:app --host 127.0.0.1 --port 8787 --workers 1
Then set up cloud execution with Modal and your LiteLLM gateway, and connect your apps. Star the repo on GitHub, open an issue, or send us a PR. Moyai will probably review it


