LiteLLM is moving to Rust Read the latest updates.
1.105.0rc3 - Decisions API
Deploy this version
- Docker
- Pip
docker run \
-e LITELLM_MASTER_KEY=sk-<paste-a-long-random-key> \
-e DATABASE_URL=postgresql://<user>:<password>@<host>:5432/<dbname> \
-e STORE_MODEL_IN_DB=True \
-p 4000:4000 \
docker.litellm.ai/berriai/litellm:1.105.0-rc.3
pip install litellm==1.105.0rc3
The published GitHub tag is v1.105.0-rc.3. These notes compare it with v1.105.0-rc.2, the previous release candidate on rc/1.105.0. It adds the Decisions API routes and providers. There are no new database migrations, and the Decisions routes are new on this line. The v1.105.0-rc.3 tag points at c09b4d5
Decisions API
The proxy now serves decision models, which answer typed questions (yes/no predicates, choices and scores) about an input. Requests go through the usual keys, budgets, spend tracking and guardrails. See the Decisions docs
- Two request formats
POST /v1/decisionsandPOST /decisionstake OpenAI's Decisions API format:inputplus aquestionslist, answered as ananswerslist with OpenAI-styleusagePOST /v1/systemoneandPOST /systemonetake the System One format:stateplus a namedquestionsmap, answered as ananswersmap keyed by question name- Both routes go through one shared format inside the proxy, so any Decisions model can be called on either route
- Providers:
openai/(gpt-6-luna),perplexity/,typesafe/,openrouter/,cloudflare/clefandcloudflare/clef-flash, and self-hostedstrands_decider/(requiresapi_base) - OpenAI: images, multi-message input,
safety_identifier, names and refusals pass through on/v1/decisions. The provider readslitellm.api_key,litellm.openai_key,litellm.api_base,OPENAI_BASE_URLandOPENAI_API_BASEin the same order as other OpenAI calls - Validation before any upstream call: a missing
state, a malformed question, more than 128 questions, a single-option question or a body that is not JSON returns 400. Image input sent to a provider that only accepts text returns a clear 400 - Spend tracking: per-token pricing from the cost map, with cached and cache-write input tokens billed at their own rates
- Health checks:
/healthprobesmode: evaluationdeployments with a one-question Decisions call, andhealth_check_paramscan override itsstateandquestions - Guardrails: OpenAI-format responses include
guardrail_informationwhen the caller setsinclude_guardrail_response - Pass-through: an existing
pass_through_endpointsentry at/v1/decisionskeeps answering that path, while/decisionsserves the native route
New Model Support (6 new models)
| Provider | Model | Context Window | Input ($/1M tokens) | Output ($/1M tokens) |
|---|---|---|---|---|
| Perplexity | perplexity/pplx-decider-v1-27b | 262K | $0.04 | $0.00 |
| Cloudflare | cloudflare/clef | 65K | $0.24 | $0.00 |
| Cloudflare | cloudflare/clef-flash | 65K | $0.09 | $0.00 |
| Cloudflare | cloudflare/@cf/cloudflare/clef | 65K | $0.24 | $0.00 |
| Cloudflare | cloudflare/@cf/cloudflare/clef-flash | 65K | $0.09 | $0.00 |
| Strands Decider | strands_decider/strands-decider-2B-hobson-v19 | - | $0.00 (self-hosted) | $0.00 |
The existing gpt-6-luna entry now lists /v1/decisions among its supported endpoints
What's Changed
- Add a unified
/v1/decisionsendpoint for System One-compatible providers - PR #44236 - Serve the System One format at
/v1/systemoneand OpenAI's format at/v1/decisions- PR #45184 - Add OpenAI as a Decisions provider behind a shared decisions format - PR #45214
All three are backported in PR #45189
Full Changelog
https://github.com/BerriAI/litellm/compare/v1.105.0-rc.2...v1.105.0-rc.3