Skip to main content
LiteLLM is moving to Rust Read the latest updates.

v1.104.2 - Decisions API

Deploy this version​

docker run \
-e LITELLM_MASTER_KEY=sk-<paste-a-long-random-key> \
-e DATABASE_URL=postgresql://<user>:<password>@<host>:5432/<dbname> \
-e STORE_MODEL_IN_DB=True \
-p 4000:4000 \
docker.litellm.ai/berriai/litellm:1.104.2

This release is published as ghcr.io/berriai/litellm:v1.104.2. See the GitHub release and the full releases page

v1.104.2 is a patch release on top of v1.104.1. It adds the Decisions API routes and providers. There are no new database migrations or breaking changes: the Decisions routes are new on this line. The v1.104.2 tag points at fc3920b

Decisions API​

The proxy now serves decision models, which answer typed questions (yes/no predicates, choices and scores) about an input. Requests go through the usual keys, budgets, spend tracking and guardrails. See the Decisions docs

  • Two request formats
    • POST /v1/decisions and POST /decisions take OpenAI's Decisions API format: input plus a questions list, answered as an answers list with OpenAI-style usage
    • POST /v1/systemone and POST /systemone take the System One format: state plus a named questions map, answered as an answers map keyed by question name
    • Both routes go through one shared format inside the proxy, so any Decisions model can be called on either route
  • Providers: openai/ (gpt-6-luna), perplexity/, typesafe/, openrouter/, cloudflare/clef and cloudflare/clef-flash, and self-hosted strands_decider/ (requires api_base)
  • OpenAI: images, multi-message input, safety_identifier, names and refusals pass through on /v1/decisions. The provider reads litellm.api_key, litellm.openai_key, litellm.api_base, OPENAI_BASE_URL and OPENAI_API_BASE in the same order as other OpenAI calls
  • Validation before any upstream call: a missing state, a malformed question, more than 128 questions, a single-option question or a body that is not JSON returns 400. Image input sent to a provider that only accepts text returns a clear 400
  • Spend tracking: per-token pricing from the cost map, with cached and cache-write input tokens billed at their own rates
  • Health checks: /health probes mode: evaluation deployments with a one-question Decisions call, and health_check_params can override its state and questions
  • Guardrails: OpenAI-format responses include guardrail_information when the caller sets include_guardrail_response
  • Pass-through: an existing pass_through_endpoints entry at /v1/decisions keeps answering that path, while /decisions serves the native route

New Model Support (6 new models)​

ProviderModelContext WindowInput ($/1M tokens)Output ($/1M tokens)
Perplexityperplexity/pplx-decider-v1-27b262K$0.04$0.00
Cloudflarecloudflare/clef65K$0.24$0.00
Cloudflarecloudflare/clef-flash65K$0.09$0.00
Cloudflarecloudflare/@cf/cloudflare/clef65K$0.24$0.00
Cloudflarecloudflare/@cf/cloudflare/clef-flash65K$0.09$0.00
Strands Deciderstrands_decider/strands-decider-2B-hobson-v19-$0.00 (self-hosted)$0.00

The existing gpt-6-luna entry now lists /v1/decisions among its supported endpoints

What's Changed​

  • Add a unified /v1/decisions endpoint for System One-compatible providers - PR #44236
  • Serve the System One format at /v1/systemone and OpenAI's format at /v1/decisions - PR #45184
  • Add OpenAI as a Decisions provider behind a shared decisions format - PR #45214

All three are backported in PR #45190

Full Changelog​

https://github.com/BerriAI/litellm/compare/v1.104.1...v1.104.2

LiteLLM Enterprise
SSO/SAML, audit logs, spend tracking, multi-team management, and guardrails, built for production.
Learn more →