Skip to main content
LiteLLM is moving to Rust Read the latest updates.

1.104.0rc1 - Claude Opus 5.5, GPT-6, Master Key Enforcement & Team Routing Controls

Deploy this version​

docker run \
-e LITELLM_MASTER_KEY=sk-<paste-a-long-random-key> \
-e DATABASE_URL=postgresql://<user>:<password>@<host>:5432/<dbname> \
-e STORE_MODEL_IN_DB=True \
-p 4000:4000 \
docker.litellm.ai/berriai/litellm:1.104.0-rc.1

The published GitHub tag is v1.104.0-rc.1. These notes compare it with v1.103.0-rc.1, the previous release candidate cut from main. Changes backported onto rc/1.103.0 and already shipped in v1.103.0 are omitted, and so are the three Usage page follow-ups built on the top-N key cap, which were reverted on rc/1.104.0 before this tag so the Usage pages load every key the same way v1.103.0 does

Customer-facing changes come first. Test, CI and internal changes are listed at the bottom

Breaking Changes

These callouts cover user-facing behavior that differs from v1.103.0, the latest stable release

The proxy refuses to start with an unset, empty, or publicly known master key. A deployment with no LITELLM_MASTER_KEY, an empty one, or sk-1234 stops booting after the upgrade. The startup error names where the bad key came from and prints a command that generates a secure one. If the database holds values encrypted with the old key, also set LITELLM_MIGRATE_FROM_MASTER_KEY so the next boot re-encrypts them. To keep the old behavior on a local sandbox, set LITELLM_DANGEROUSLY_PERMIT_WEAK_OR_UNSET_MASTER_KEY=true or general_settings.dangerously_permit_weak_or_unset_master_key: true. See PR #42019, PR #42011

An exhausted budget now returns HTTP 422 instead of 429. Clients stop treating a spent budget as a retryable rate limit. Real rpm/tpm limits still return 429. Set litellm_settings.budget_exceeded_status_code: 429 to keep the old status. See PR #42097

Key Highlights​

  • New frontier models on day one: Claude Opus 5.5 across Anthropic, Bedrock, Vertex AI and Azure AI, and GPT-6 Sol and GPT-6 Luna across OpenAI, Bedrock and Azure Foundry, among 331 new catalog entries
  • New providers and routes: Eden AI and Nadir providers, TinyFish, fal.ai queue and OpenRouter decisions pass-through routes, OpenAI models on Bedrock's native Responses API, and Claude on Bedrock Mantle's native Messages API
  • Gateway hardening: the proxy refuses a weak or missing master key, dashboard sign-in adds breached-password detection and forced password resets, sessions are revoked on logout, and auth fails closed during a database outage
  • Routing controls: group-scoped priority routing, time-windowed team reservation of deployments, native compact-to-fit across conversation APIs, a JEV classifier for the Auto Router, and configurable provider affinity headers
  • Admin UI: the LiteAdmin assistant, prompt caching savings, internal-user savings and Auto Router usage, and Capability and Fuse v2 routing forecasts

New Providers and Endpoints​

New Providers (2 new providers)​

ProviderSupported LiteLLM EndpointsDescription
Eden AI/v1/chat/completions, /v1/responses, /v1/messages, /v1/embeddings, audio, images, /v1/videosEden AI's unified API across chat, embeddings, audio, image and video models
Nadir/v1/chat/completionsNadir's nadir/auto router, with the cost Nadir reports recorded as the response cost

Expanded provider endpoint support​

ProviderEndpointWhat you can do
TinyFish/tinyfish/*Run TinyFish Agent API automations through a pass-through route with per-step billing
fal.ai/fal_ai/*Submit and poll fal queue jobs through a pass-through route with spend tracking
OpenRouter/openrouter/alpha/decisionsReach OpenRouter's decisions API through a pass-through route
Amazon Bedrock/v1/responsesServe OpenAI models on bedrock-runtime's native Responses API
Bedrock Mantle/v1/messagesServe Claude models on Mantle's native Anthropic Messages API
Vertex AI/v1/files, /v1/batchesSubmit native Vertex batch JSONL with cost tracking

New Models / Updated Models​

New Model Support (331 new models)​

Counts represent new catalog identifiers, including aliases and regional variants. Prices below are the values bundled in this release, in USD; runtime pricing-map reloads can update them. Input and output columns show base token rates; long-context, cache, image-token, and other specialized rates depend on the model

ProviderModelContext WindowInput ($/1M tokens)Output ($/1M tokens)Features / special pricing
Amazon Bedrockanthropic.claude-mythos-5-11,000,000$10$50Chat; Reasoning; Vision; Tool calling; Prompt caching; Structured output; PDF input; Computer use
Amazon Bedrockanthropic.claude-opus-5-51,000,000$4$20Chat; Reasoning; Vision; Tool calling; Prompt caching; Structured output; PDF input; Computer use
Amazon Bedrockapac.anthropic.claude-fable-51,000,000$11$55Chat; Reasoning; Vision; Tool calling; Prompt caching; Structured output; PDF input; Computer use
Amazon Bedrockapac.anthropic.claude-mythos-preview1,000,000$27.5$137.5Chat; Reasoning; Vision; Tool calling; Prompt caching
Amazon Bedrockapac.anthropic.claude-opus-4-71,000,000$5.5$27.5Chat; Reasoning; Vision; Tool calling; Prompt caching; Structured output; PDF input; Computer use
Amazon Bedrockapac.anthropic.claude-opus-4-81,000,000$5.5$27.5Chat; Reasoning; Vision; Tool calling; Prompt caching; Structured output; PDF input; Computer use
Amazon Bedrockapac.anthropic.claude-opus-51,000,000$5.5$27.5Chat; Reasoning; Vision; Tool calling; Prompt caching; Structured output; PDF input; Computer use
Amazon Bedrockapac.anthropic.claude-opus-5-51,000,000$4.4$22Chat; Reasoning; Vision; Tool calling; Prompt caching; Structured output; PDF input; Computer use
Amazon Bedrockapac.anthropic.claude-sonnet-4-61,000,000$3.3$16.5Chat; Reasoning; Vision; Tool calling; Prompt caching; Structured output; PDF input; Computer use
Amazon Bedrockapac.anthropic.claude-sonnet-51,000,000$2.2$11Chat; Reasoning; Vision; Tool calling; Prompt caching; Structured output; PDF input; Computer use
Amazon Bedrockau.anthropic.claude-fable-51,000,000$11$55Chat; Reasoning; Vision; Tool calling; Prompt caching; Structured output; PDF input; Computer use
Amazon Bedrockau.anthropic.claude-mythos-preview1,000,000$27.5$137.5Chat; Reasoning; Vision; Tool calling; Prompt caching
Amazon Bedrockau.anthropic.claude-opus-5-51,000,000$4.4$22Chat; Reasoning; Vision; Tool calling; Prompt caching; Structured output; PDF input; Computer use
Amazon Bedrockbedrock/ap-northeast-1/qwen.qwen3-next-80b-a3b128,000$0.18$1.45Chat; Tool calling; Structured output
Amazon Bedrockbedrock/ap-south-1/qwen.qwen3-next-80b-a3b128,000$0.18$1.41Chat; Tool calling; Structured output
Amazon Bedrockbedrock/ap-southeast-2/qwen.qwen3-next-80b-a3b128,000$0.1545$1.236Chat; Tool calling; Structured output
Amazon Bedrockbedrock/eu-west-1/qwen.qwen3-next-80b-a3b128,000$0.18$1.41Chat; Tool calling; Structured output
Amazon Bedrockbedrock/eu-west-2/nvidia.nemotron-super-3-120b256,000$0.23$1.01Chat; Reasoning; Tool calling; Structured output
Amazon Bedrockbedrock/eu-west-2/qwen.qwen3-next-80b-a3b128,000$0.23$1.86Chat; Tool calling; Structured output
Amazon Bedrockbedrock/sa-east-1/qwen.qwen3-next-80b-a3b128,000$0.18$1.45Chat; Tool calling; Structured output
Amazon Bedrockbedrock/us-gov-east-1/anthropic.claude-opus-5-51,000,000$4.8$24Chat; Reasoning; Vision; Tool calling; Prompt caching; Structured output; PDF input; Computer use
Amazon Bedrockbedrock/us-gov-west-1/anthropic.claude-opus-5-51,000,000$4.8$24Chat; Reasoning; Vision; Tool calling; Prompt caching; Structured output; PDF input; Computer use
Amazon Bedrockdeepseek.r1-v1:0128,000$1.35$5.4Chat; Reasoning
Amazon Bedrockeu.anthropic.claude-opus-5-51,000,000$4.4$22Chat; Reasoning; Vision; Tool calling; Prompt caching; Structured output; PDF input; Computer use
Amazon Bedrockglobal.anthropic.claude-mythos-5-11,000,000$10$50Chat; Reasoning; Vision; Tool calling; Prompt caching; Structured output; PDF input; Computer use
Amazon Bedrockglobal.anthropic.claude-opus-5-51,000,000$4$20Chat; Reasoning; Vision; Tool calling; Prompt caching; Structured output; PDF input; Computer use
Amazon Bedrockglobal.moonshotai.kimi-k31,000,000$3$15Chat; Reasoning; Vision; Tool calling; Prompt caching; Structured output
Amazon Bedrockglobal.openai.gpt-5.41,000,000$2.5$15Chat; Reasoning; Vision; Tool calling
Amazon Bedrockglobal.openai.gpt-5.51,000,000$5$30Chat; Reasoning; Vision; Tool calling
Amazon Bedrockglobal.openai.gpt-6-luna1,050,000$0.1$0.5Chat; Reasoning; Vision; Tool calling; Prompt caching
Amazon Bedrockglobal.openai.gpt-6-sol1,050,000$2$10Chat; Reasoning; Vision; Tool calling; Prompt caching
Amazon Bedrockjp.anthropic.claude-opus-5-51,000,000$4.4$22Chat; Reasoning; Vision; Tool calling; Prompt caching; Structured output; PDF input; Computer use
Amazon Bedrockmistral.pixtral-large-2502-v1:0128,000$2$6Chat; Tool calling
Amazon Bedrockmoonshotai.kimi-k31,000,000$3$15Chat; Reasoning; Vision; Tool calling; Prompt caching; Structured output
Amazon Bedrockopenai.gpt-6-luna1,050,000$0.1$0.5Chat; Reasoning; Vision; Tool calling; Prompt caching
Amazon Bedrockopenai.gpt-6-sol1,050,000$2$10Chat; Reasoning; Vision; Tool calling; Prompt caching
Amazon Bedrockus-gov.anthropic.claude-opus-5-51,000,000$4.8$24Chat; Reasoning; Vision; Tool calling; Prompt caching; Structured output; PDF input; Computer use
Amazon Bedrockus.anthropic.claude-mythos-51,000,000$11$55Chat; Reasoning; Vision; Tool calling; Prompt caching; Structured output; PDF input; Computer use
Amazon Bedrockus.anthropic.claude-mythos-5-11,000,000$11$55Chat; Reasoning; Vision; Tool calling; Prompt caching; Structured output; PDF input; Computer use
Amazon Bedrockus.anthropic.claude-mythos-preview1,000,000$27.5$137.5Chat; Reasoning; Vision; Tool calling; Prompt caching
Amazon Bedrockus.anthropic.claude-opus-5-51,000,000$4.4$22Chat; Reasoning; Vision; Tool calling; Prompt caching; Structured output; PDF input; Computer use
Amazon Bedrockus.moonshotai.kimi-k31,000,000$3.3$16.5Chat; Reasoning; Vision; Tool calling; Prompt caching; Structured output
Amazon Bedrockus.openai.gpt-5.41,000,000$2.75$16.5Chat; Reasoning; Vision; Tool calling
Amazon Bedrockus.openai.gpt-5.51,000,000$5.5$33Chat; Reasoning; Vision; Tool calling
Amazon Bedrockus.openai.gpt-6-luna1,050,000$0.11$0.55Chat; Reasoning; Vision; Tool calling; Prompt caching
Amazon Bedrockus.openai.gpt-6-sol1,050,000$2.2$11Chat; Reasoning; Vision; Tool calling; Prompt caching
Anthropicclaude-opus-5-51,000,000$4$20Chat; Reasoning; Vision; Tool calling; Prompt caching; Structured output; PDF input; Computer use
Azure AIazure_ai/Cohere-command-a-plus-05-2026128,000$0.8$3.2Chat; Reasoning; Tool calling
Azure AIazure_ai/FW-DeepSeek-V4-Flash1,000,000$0.15$0.31Chat; Reasoning; Tool calling; Prompt caching
Azure AIazure_ai/FW-DeepSeek-V4.1-Flash1,000,000$0.375$1.5Chat; Reasoning; Tool calling; Prompt caching
Azure AIazure_ai/FW-GLM-5.31,048,576$1.75$5.5Chat; Reasoning; Tool calling; Prompt caching
Azure AIazure_ai/FW-GLM-5.3-Flash1,048,576$0.188$0.625Chat; Reasoning; Tool calling; Prompt caching
Azure AIazure_ai/FW-GPT-OSS-120B131,072$0.165$0.66Chat; Reasoning; Tool calling; Prompt caching; Structured output
Azure AIazure_ai/MAI-Image-2.5-Pro-$5-Image generation
Azure AIazure_ai/MAI-Image-2.6-$5-Image generation
Azure AIazure_ai/MAI-Image-2.6-Flash-$1.75-Image generation
Azure AIazure_ai/claude-opus-5-51,000,000$4$20Chat; Reasoning; Vision; Tool calling; Prompt caching; Structured output; PDF input; Computer use
Azure AIazure_ai/deepseek-v4.1-flash1,000,000$0.375$1.5Chat; Reasoning; Tool calling; Prompt caching
Azure AIazure_ai/gpt-6-luna922,000$0.1$0.5Chat; Reasoning; Vision; Tool calling; Prompt caching; Structured output; PDF input; Web search; Computer use
Azure AIazure_ai/gpt-6-sol922,000$2$10Chat; Reasoning; Vision; Tool calling; Prompt caching; Structured output; PDF input; Web search; Computer use
Azure AIazure_ai/gpt-image-2-$5-Image generation; Vision
Azure AIazure_ai/mistral-medium-3-5128,000$1.5$7.5Chat; Vision; Structured output
Azure AIazure_ai/muse-spark-1.31,048,576$1.25$4.25Chat; Reasoning; Vision; Tool calling; Prompt caching; Structured output
Azure OpenAIazure/eu/gpt-6-luna922,000$0.12$0.6Chat; Reasoning; Vision; Tool calling; Prompt caching; Structured output; PDF input; Web search; Computer use
Azure OpenAIazure/eu/gpt-6-sol922,000$2.4$12Chat; Reasoning; Vision; Tool calling; Prompt caching; Structured output; PDF input; Web search; Computer use
Azure OpenAIazure/gpt-6-luna922,000$0.1$0.5Chat; Reasoning; Vision; Tool calling; Prompt caching; Structured output; PDF input; Web search; Computer use
Azure OpenAIazure/gpt-6-luna-2026-09-22922,000$0.1$0.5Chat; Reasoning; Vision; Tool calling; Prompt caching; Structured output; PDF input; Web search; Computer use
Azure OpenAIazure/gpt-6-sol922,000$2$10Chat; Reasoning; Vision; Tool calling; Prompt caching; Structured output; PDF input; Web search; Computer use
Azure OpenAIazure/gpt-6-sol-2026-09-22922,000$2$10Chat; Reasoning; Vision; Tool calling; Prompt caching; Structured output; PDF input; Web search; Computer use
Azure OpenAIazure/gpt-audio128,000$2.5$10Chat; Tool calling
Azure OpenAIazure/gpt-audio-1.5128,000$2.5$10Chat; Tool calling
Azure OpenAIazure/gpt-live-1---Realtime; Tool calling; Audio input; input_cost_per_second: $0.00083333
Azure OpenAIazure/gpt-live-transcribe32,000--Transcription; Audio input; input_cost_per_second: $0.00028333
Azure OpenAIazure/gpt-realtime32,000$4$16Realtime; Tool calling; Audio input
Azure OpenAIazure/gpt-realtime-1.532,000$4$16Realtime; Tool calling; Audio input
Azure OpenAIazure/gpt-realtime-translate32,000--Realtime; Audio input; input_cost_per_second: $0.00056667
Azure OpenAIazure/gpt-transcribe---Transcription; Audio input; input_cost_per_second: $0.000075
Azure OpenAIazure/us/gpt-6-luna922,000$0.11$0.55Chat; Reasoning; Vision; Tool calling; Prompt caching; Structured output; PDF input; Web search; Computer use
Azure OpenAIazure/us/gpt-6-sol922,000$2.2$11Chat; Reasoning; Vision; Tool calling; Prompt caching; Structured output; PDF input; Web search; Computer use
Basetenbaseten/deepseek-ai/DeepSeek-V4-Flash-07311,048,576$0.13$0.26Chat; Reasoning; Tool calling; Prompt caching; Structured output
Basetenbaseten/deepseek-ai/DeepSeek-V4-Pro1,048,576$1.74$3.48Chat; Reasoning; Tool calling; Prompt caching; Structured output
Basetenbaseten/deepseek-ai/DeepSeek-V4-Pro-08131,048,576$1.32$3.96Chat; Reasoning; Tool calling; Prompt caching; Structured output
Basetenbaseten/deepseek-ai/DeepSeek-V4.1-Flash1,048,576$0.3$1.2Chat; Reasoning; Vision; Tool calling; Prompt caching; Structured output
Basetenbaseten/moonshotai/Kimi-K2.6262,000$0.95$4Chat; Reasoning; Vision; Tool calling; Prompt caching; Structured output
Basetenbaseten/moonshotai/Kimi-K2.7-Code262,000$0.95$4Chat; Reasoning; Vision; Tool calling; Prompt caching; Structured output
Basetenbaseten/moonshotai/Kimi-K31,048,576$3$15Chat; Reasoning; Vision; Tool calling; Prompt caching; Structured output
Basetenbaseten/nvidia/NVIDIA-Nemotron-3-Ultra-550B-A55B202,800$0.6$2.4Chat; Reasoning; Tool calling; Prompt caching; Structured output
Basetenbaseten/thinkingmachines/inkling1,048,576$1$4.05Chat; Reasoning; Vision; Tool calling; Prompt caching; Structured output
Basetenbaseten/thinkingmachines/inkling-small1,048,576$0.5$1.2Chat; Reasoning; Vision; Tool calling; Prompt caching; Structured output
Basetenbaseten/zai-org/GLM-5.21,048,576$1.4$4.4Chat; Reasoning; Vision; Tool calling; Prompt caching; Structured output
Basetenbaseten/zai-org/GLM-5.2-Fast1,048,576$2.1$6.6Chat; Reasoning; Vision; Tool calling; Prompt caching; Structured output
Basetenbaseten/zai-org/GLM-5.3-Fast1,048,576$2.1$6.6Chat; Reasoning; Vision; Tool calling; Prompt caching; Structured output
Basetenbaseten/zai-org/GLM-5.3-Flash1,048,576$0.15$0.5Chat; Reasoning; Vision; Tool calling; Prompt caching; Structured output
Bedrock Mantlebedrock_mantle/anthropic.claude-haiku-4-5200,000$1$5Chat; Reasoning; Vision; Tool calling; Prompt caching; Structured output; PDF input; Computer use
Bedrock Mantlebedrock_mantle/deepseek.v3.1128,000$0.58$1.68Chat; Tool calling
Bedrock Mantlebedrock_mantle/moonshotai.kimi-k2-thinking256,000$0.6$2.5Chat; Reasoning; Tool calling
Bedrock Mantlebedrock_mantle/openai.gpt-6-luna1,050,000$0.11$0.55Responses; Reasoning; Vision; Tool calling; Prompt caching; Structured output
Bedrock Mantlebedrock_mantle/openai.gpt-6-sol1,050,000$2.2$11Responses; Reasoning; Vision; Tool calling; Prompt caching; Structured output
Bedrock Mantlebedrock_mantle/qwen.qwen3-235b-a22b-2507256,000$0.22$0.88Chat; Reasoning; Tool calling
Bedrock Mantlebedrock_mantle/qwen.qwen3-32b32,000$0.15$0.6Chat; Reasoning; Tool calling
Bedrock Mantlebedrock_mantle/qwen.qwen3-coder-30b-a3b-instruct256,000$0.15$0.6Chat; Tool calling
Bedrock Mantlebedrock_mantle/qwen.qwen3-coder-480b-a35b-instruct128,000$0.45$1.8Chat; Tool calling
Bedrock Mantlebedrock_mantle/qwen.qwen3-next-80b-a3b-instruct256,000$0.14$1.2Chat; Reasoning; Tool calling
Bedrock Mantlebedrock_mantle/qwen.qwen3-vl-235b-a22b-instruct256,000$0.53$2.66Chat; Vision; Tool calling
Coherec4ai-aya-expanse-32b128,000$0.5$1.5Chat
Fireworks AIfireworks_ai/accounts/fireworks/models/ember-11,048,576$3$15Chat; Reasoning; Vision; Tool calling; Structured output
Fireworks AIfireworks_ai/accounts/fireworks/routers/deepseek-v4p1-flash-us1,048,576$0.45$1.8Chat; Reasoning; Vision; Tool calling; Prompt caching; Structured output
Fireworks AIfireworks_ai/accounts/fireworks/routers/glm-5p3-flash-us1,048,576$0.225$0.75Chat; Vision; Tool calling; Structured output
Fireworks AIfireworks_ai/accounts/fireworks/routers/glm-5p3-us1,048,576$2.1$6.6Chat; Reasoning; Tool calling; Structured output
Fireworks AIfireworks_ai/deepseek-v4-pro-08131,048,576$1.32$3.96Chat; Reasoning; Tool calling; Structured output
Fireworks AIfireworks_ai/deepseek-v4p1-flash-us1,048,576$0.45$1.8Chat; Reasoning; Vision; Tool calling; Prompt caching; Structured output
Fireworks AIfireworks_ai/glm-5p31,048,576$1.4$4.4Chat; Reasoning; Tool calling; Structured output
Fireworks AIfireworks_ai/glm-5p3-flash1,048,576$0.15$0.5Chat; Vision; Tool calling; Structured output
Fireworks AIfireworks_ai/glm-5p3-flash-us1,048,576$0.225$0.75Chat; Vision; Tool calling; Structured output
Fireworks AIfireworks_ai/glm-5p3-us1,048,576$2.1$6.6Chat; Reasoning; Tool calling; Structured output
Geminigemini/deep-research-max-preview-04-2026131,072$2$12Chat; Vision; Web search
Geminigemini/deep-research-preview-04-2026131,072$2$12Chat; Vision; Web search
Geminigemini/gemini-3.8-flash-lite-tts8,192$0.5$6Speech
Geminigemini/gemini-3.8-flash-tts8,192$0.5$9Speech
Geminigemini/lyria-realtime-exp1,048,576$0$0Chat
Groqgroq/llama-guard-3-8b8,192$0.2$0.2Chat
Nebiusnebius/deepseek-ai/DeepSeek-V4.1-Flash1,048,576$0.3$1.2Chat; Vision
OpenAIgpt-6-luna922,000$0.1$0.5Chat; Reasoning; Vision; Tool calling; Prompt caching; Structured output; PDF input; Web search; Computer use
OpenAIgpt-6-sol922,000$2$10Chat; Reasoning; Vision; Tool calling; Prompt caching; Structured output; PDF input; Web search; Computer use
OpenRouteropenrouter/aion-labs/aion-3.5262,144$3$6Chat; Reasoning; Tool calling; Prompt caching; Structured output
OpenRouteropenrouter/aion-labs/aion-3.5-mini262,144$0.7$1.4Chat; Reasoning; Tool calling; Prompt caching; Structured output
OpenRouteropenrouter/anthropic/claude-opus-5.51,000,000$4$20Chat; Reasoning; Vision; Tool calling; Prompt caching; Structured output; PDF input; Web search
OpenRouteropenrouter/anthropic/claude-opus-5.5:batch1,000,000$2$10Chat; Reasoning; Vision; Tool calling; Prompt caching; Structured output; PDF input; Web search
OpenRouteropenrouter/cohere/command-a-plus192,000$0.3$1.5Chat; Reasoning; Vision; Tool calling; Prompt caching; Structured output
OpenRouteropenrouter/deepseek/deepseek-v4.1-flash:batch1,048,576$0.112$0.336Chat; Reasoning; Vision; Tool calling; Prompt caching; Structured output
OpenRouteropenrouter/fireworks/ember-11,048,576$3$15Chat; Reasoning; Vision; Tool calling; Prompt caching; Structured output
OpenRouteropenrouter/nex-agi/nex-n2.5-mini262,144$0.025$0.1Chat; Reasoning; Vision; Prompt caching; Structured output
OpenRouteropenrouter/nex-agi/nex-n2.5-pro262,144$0.075$0.25Chat; Reasoning; Vision; Tool calling; Prompt caching; Structured output
OpenRouteropenrouter/openai/gpt-6-luna1,050,000$0.1$0.5Chat; Reasoning; Vision; Tool calling; Prompt caching; Structured output; PDF input; Web search
OpenRouteropenrouter/openai/gpt-6-luna-pro1,050,000$0.1$0.5Chat; Reasoning; Vision; Tool calling; Prompt caching; Structured output; PDF input; Web search
OpenRouteropenrouter/openai/gpt-6-luna-pro:batch1,050,000$0.05$0.25Chat; Reasoning; Vision; Tool calling; Prompt caching; Structured output; PDF input; Web search
OpenRouteropenrouter/openai/gpt-6-luna:batch1,050,000$0.05$0.25Chat; Reasoning; Vision; Tool calling; Prompt caching; Structured output; PDF input; Web search
OpenRouteropenrouter/openai/gpt-6-sol1,050,000$2$10Chat; Reasoning; Vision; Tool calling; Prompt caching; Structured output; PDF input; Web search
OpenRouteropenrouter/openai/gpt-6-sol-pro1,050,000$2$10Chat; Reasoning; Vision; Tool calling; Prompt caching; Structured output; PDF input; Web search
OpenRouteropenrouter/openai/gpt-6-sol-pro:batch1,050,000$1$5Chat; Reasoning; Vision; Tool calling; Prompt caching; Structured output; PDF input; Web search
OpenRouteropenrouter/openai/gpt-6-sol:batch1,050,000$1$5Chat; Reasoning; Vision; Tool calling; Prompt caching; Structured output; PDF input; Web search
OpenRouteropenrouter/openai/gpt-oss-20b:batch131,072$0.024$0.112Chat; Reasoning; Tool calling; Structured output
OpenRouteropenrouter/perceptron/perceptron-mk1.536,864$0.15$1.5Chat; Reasoning; Vision; Tool calling; Structured output; Audio input; Video input
OpenRouteropenrouter/qwen/qwen3.8-max-prime1,000,000$4$12Chat; Reasoning; Vision; Tool calling; Prompt caching; Structured output; Video input
OpenRouteropenrouter/qwen/qwen3.8-omni-flash1,000,000$0.15$0.47Chat; Reasoning; Vision; Tool calling; Prompt caching; Structured output; Audio input
OpenRouteropenrouter/stealth/space-bunny-alpha1,000,000$0$0Chat; Reasoning; Vision; Tool calling; Video input
OpenRouteropenrouter/typesafe/jev-1.1332,000$0.042$0evaluation
OpenRouteropenrouter/upstage/solar-mini4524,288$0.05$0.2Chat; Reasoning; Tool calling; Prompt caching; Structured output
OpenRouteropenrouter/x-ai/grok-4.7500,000$1.6$4.8Chat; Reasoning; Vision; Tool calling; Prompt caching; Structured output; PDF input; Web search
OpenRouteropenrouter/xiaomi/mimo-v2.6-flash1,048,576$0.14$0.28Chat; Reasoning; Vision; Tool calling; Prompt caching; Structured output; Audio input
OpenRouteropenrouter/xiaomi/mimo-v2.6-pro1,048,576$0.435$0.87Chat; Reasoning; Vision; Tool calling; Prompt caching; Structured output; Audio input
OpenRouteropenrouter/xiaomi/mimo-v2.6-pro-ultraspeed1,048,576$4.35$8.7Chat; Reasoning; Vision; Tool calling; Prompt caching; Structured output; Audio input
OpenRouteropenrouter/z-ai/glm-5.3-prime1,000,000$2.8$8.8Chat; Reasoning; Tool calling; Prompt caching; Structured output
Together AItogether_ai/together/Tev1-4B-experimental32,768$0.042$0Chat
Vertex AIgemini-3.8-flash-cyber1,048,576$1.5$7.5Chat; Reasoning; Vision; Prompt caching; Structured output; PDF input; Audio input; Video input
Vertex AIvertex_ai/chirp_2---Transcription; input_cost_per_second: $0.00026667
Vertex AIvertex_ai/claude-opus-5-51,000,000$4$20Chat; Reasoning; Vision; Tool calling; Prompt caching; Structured output; PDF input; Computer use
Vertex AIvertex_ai/claude-opus-5-5@default1,000,000$4$20Chat; Reasoning; Vision; Tool calling; Prompt caching; Structured output; PDF input; Computer use
Vertex AIvertex_ai/gemini-2.0-flash-$0.15$0.6Chat
Vertex AIvertex_ai/gemini-2.0-flash-lite-$0.075$0.3Chat
Vertex AIvertex_ai/gemini-2.5-flash-tts-$0.5$10Speech
Vertex AIvertex_ai/gemini-2.5-pro-tts-$1$20Speech
Vertex AIvertex_ai/gemini-3.8-flash-cyber1,048,576$1.5$7.5Chat; Reasoning; Vision; Prompt caching; Structured output; PDF input; Audio input; Video input
Vertex AIvertex_ai/gemini-3.8-live131,072$0.75$4.5Realtime; Vision; Tool calling; Web search; Audio input
Vertex AIvertex_ai/gemini-omni-1.1-flash-preview57,920$1.5$9Chat; Reasoning; Vision; Video input
Vertex AIvertex_ai/meta/llama-3.3-70b-instruct-maas128,000$0.72$0.72Chat; Tool calling
Vertex AIvertex_ai/virtual-try-on-001---Image generation; output_cost_per_image: $0.06
Vertex AIvertex_ai/zai-org/glm-5.2-maas1,000,000$1.4$4.4Chat; Reasoning; Tool calling; Prompt caching; Structured output
Weights & Biaseswandb/deepseek-ai/DeepSeek-V4.1-Flash1,049,000$0.2$0.65Chat; Reasoning; Vision; Prompt caching
Weights & Biaseswandb/google/gemma-4-26B-A4B-it262,000$0.1$0.3Chat; Reasoning; Vision; Tool calling; Prompt caching
Xiaomi MiMoxiaomi_mimo/mimo-v2.6-flash1,048,576$0.14$0.28Chat; Reasoning; Vision; Tool calling; Prompt caching; Structured output; Audio input; Video input
Xiaomi MiMoxiaomi_mimo/mimo-v2.6-pro1,048,576$0.435$0.87Chat; Reasoning; Vision; Tool calling; Prompt caching; Structured output; Audio input; Video input
fal.aifal_ai/bytedance/seedance-2.0/image-to-video---Video generation; output_cost_per_second: $0.3034; output_cost_per_second_480p: $0.1346; output_cost_per_second_720p: $0.3034
fal.aifal_ai/bytedance/seedance-2.0/reference-to-video---Video generation; output_cost_per_second: $0.3034; output_cost_per_second_480p: $0.1346; output_cost_per_second_720p: $0.3034
fal.aifal_ai/bytedance/seedance-2.0/text-to-video---Video generation; output_cost_per_second: $0.3034; output_cost_per_second_480p: $0.1346; output_cost_per_second_720p: $0.3034
fal.aifal_ai/bytedance/seedance-2.5/image-to-video---Video generation; output_cost_per_second: $0.473; output_cost_per_second_480p: $0.2205; output_cost_per_second_720p: $0.473
fal.aifal_ai/bytedance/seedance-2.5/reference-to-video---Video generation; output_cost_per_second: $0.473; output_cost_per_second_480p: $0.2205; output_cost_per_second_720p: $0.473
fal.aifal_ai/bytedance/seedance-2.5/text-to-video---Video generation; output_cost_per_second: $0.473; output_cost_per_second_480p: $0.2205; output_cost_per_second_720p: $0.473
fal.aifal_ai/fal-ai/flux-lora-depth---Image generation; output_cost_per_image: $0.035; output_cost_per_pixel: $0.00000003
fal.aifal_ai/fal-ai/flux/dev---Image generation; output_cost_per_image: $0.025; output_cost_per_pixel: $0.00000002
fal.aifal_ai/fal-ai/moondream3-preview/query-$0.4$3.5Chat; Reasoning; Vision
fal.aifal_ai/fal-ai/nano-banana-2---Image generation; output_cost_per_image: $0.08; output_cost_per_image_0.5K: $0.06; output_cost_per_image_1K: $0.08
fal.aifal_ai/fal-ai/nano-banana-pro---Image generation; output_cost_per_image: $0.15; output_cost_per_image_1K: $0.15; output_cost_per_image_2K: $0.15
fal.aifal_ai/fal-ai/trellis---Image generation; output_cost_per_image: $0.02
fal.aifal_ai/fal-ai/trellis-2---Image generation; output_cost_per_image: $0.3; output_cost_per_image_512: $0.25; output_cost_per_image_1024: $0.3
fal.aifal_ai/high/1024-x-1024/openai/gpt-image-2.5/flare/edit---Image generation; Vision; output_cost_per_image: $0.05268
fal.aifal_ai/high/1024-x-1024/openai/gpt-image-2.5/flare/text-to-image---Image generation; Vision; output_cost_per_image: $0.05268
fal.aifal_ai/high/1024-x-1024/openai/gpt-image-2.5/sunburst/edit---Image generation; Vision; output_cost_per_image: $0.05268
fal.aifal_ai/high/1024-x-1024/openai/gpt-image-2.5/sunburst/text-to-image---Image generation; Vision; output_cost_per_image: $0.05268
fal.aifal_ai/high/1024-x-1536/openai/gpt-image-2.5/flare/edit---Image generation; Vision; output_cost_per_image: $0.04116
fal.aifal_ai/high/1024-x-1536/openai/gpt-image-2.5/flare/text-to-image---Image generation; Vision; output_cost_per_image: $0.04116
fal.aifal_ai/high/1024-x-1536/openai/gpt-image-2.5/sunburst/edit---Image generation; Vision; output_cost_per_image: $0.04116
fal.aifal_ai/high/1024-x-1536/openai/gpt-image-2.5/sunburst/text-to-image---Image generation; Vision; output_cost_per_image: $0.04116
fal.aifal_ai/high/1024-x-768/openai/gpt-image-2.5/flare/edit---Image generation; Vision; output_cost_per_image: $0.03612
fal.aifal_ai/high/1024-x-768/openai/gpt-image-2.5/flare/text-to-image---Image generation; Vision; output_cost_per_image: $0.03612
fal.aifal_ai/high/1024-x-768/openai/gpt-image-2.5/sunburst/edit---Image generation; Vision; output_cost_per_image: $0.03612
fal.aifal_ai/high/1024-x-768/openai/gpt-image-2.5/sunburst/text-to-image---Image generation; Vision; output_cost_per_image: $0.03612
fal.aifal_ai/high/1920-x-1080/openai/gpt-image-2.5/flare/edit---Image generation; Vision; output_cost_per_image: $0.0396
fal.aifal_ai/high/1920-x-1080/openai/gpt-image-2.5/flare/text-to-image---Image generation; Vision; output_cost_per_image: $0.0396
fal.aifal_ai/high/1920-x-1080/openai/gpt-image-2.5/sunburst/edit---Image generation; Vision; output_cost_per_image: $0.0396
fal.aifal_ai/high/1920-x-1080/openai/gpt-image-2.5/sunburst/text-to-image---Image generation; Vision; output_cost_per_image: $0.0396
fal.aifal_ai/high/2560-x-1440/openai/gpt-image-2.5/flare/edit---Image generation; Vision; output_cost_per_image: $0.05529
fal.aifal_ai/high/2560-x-1440/openai/gpt-image-2.5/flare/text-to-image---Image generation; Vision; output_cost_per_image: $0.05529
fal.aifal_ai/high/2560-x-1440/openai/gpt-image-2.5/sunburst/edit---Image generation; Vision; output_cost_per_image: $0.05529
fal.aifal_ai/high/2560-x-1440/openai/gpt-image-2.5/sunburst/text-to-image---Image generation; Vision; output_cost_per_image: $0.05529
fal.aifal_ai/high/3840-x-2160/openai/gpt-image-2.5/flare/edit---Image generation; Vision; output_cost_per_image: $0.10008
fal.aifal_ai/high/3840-x-2160/openai/gpt-image-2.5/flare/text-to-image---Image generation; Vision; output_cost_per_image: $0.10008
fal.aifal_ai/high/3840-x-2160/openai/gpt-image-2.5/sunburst/edit---Image generation; Vision; output_cost_per_image: $0.10008
fal.aifal_ai/high/3840-x-2160/openai/gpt-image-2.5/sunburst/text-to-image---Image generation; Vision; output_cost_per_image: $0.10008
fal.aifal_ai/low/1024-x-1024/openai/gpt-image-2.5/flare/edit---Image generation; Vision; output_cost_per_image: $0.00588
fal.aifal_ai/low/1024-x-1024/openai/gpt-image-2.5/flare/text-to-image---Image generation; Vision; output_cost_per_image: $0.00588
fal.aifal_ai/low/1024-x-1024/openai/gpt-image-2.5/sunburst/edit---Image generation; Vision; output_cost_per_image: $0.00588
fal.aifal_ai/low/1024-x-1024/openai/gpt-image-2.5/sunburst/text-to-image---Image generation; Vision; output_cost_per_image: $0.00588
fal.aifal_ai/low/1024-x-1536/openai/gpt-image-2.5/flare/edit---Image generation; Vision; output_cost_per_image: $0.00474
fal.aifal_ai/low/1024-x-1536/openai/gpt-image-2.5/flare/text-to-image---Image generation; Vision; output_cost_per_image: $0.00474
fal.aifal_ai/low/1024-x-1536/openai/gpt-image-2.5/sunburst/edit---Image generation; Vision; output_cost_per_image: $0.00474
fal.aifal_ai/low/1024-x-1536/openai/gpt-image-2.5/sunburst/text-to-image---Image generation; Vision; output_cost_per_image: $0.00474
fal.aifal_ai/low/1024-x-768/openai/gpt-image-2.5/flare/edit---Image generation; Vision; output_cost_per_image: $0.00402
fal.aifal_ai/low/1024-x-768/openai/gpt-image-2.5/flare/text-to-image---Image generation; Vision; output_cost_per_image: $0.00402
fal.aifal_ai/low/1024-x-768/openai/gpt-image-2.5/sunburst/edit---Image generation; Vision; output_cost_per_image: $0.00402
fal.aifal_ai/low/1024-x-768/openai/gpt-image-2.5/sunburst/text-to-image---Image generation; Vision; output_cost_per_image: $0.00402
fal.aifal_ai/low/1920-x-1080/openai/gpt-image-2.5/flare/edit---Image generation; Vision; output_cost_per_image: $0.00441
fal.aifal_ai/low/1920-x-1080/openai/gpt-image-2.5/flare/text-to-image---Image generation; Vision; output_cost_per_image: $0.00441
fal.aifal_ai/low/1920-x-1080/openai/gpt-image-2.5/sunburst/edit---Image generation; Vision; output_cost_per_image: $0.00441
fal.aifal_ai/low/1920-x-1080/openai/gpt-image-2.5/sunburst/text-to-image---Image generation; Vision; output_cost_per_image: $0.00441
fal.aifal_ai/low/2560-x-1440/openai/gpt-image-2.5/flare/edit---Image generation; Vision; output_cost_per_image: $0.00615
fal.aifal_ai/low/2560-x-1440/openai/gpt-image-2.5/flare/text-to-image---Image generation; Vision; output_cost_per_image: $0.00615
fal.aifal_ai/low/2560-x-1440/openai/gpt-image-2.5/sunburst/edit---Image generation; Vision; output_cost_per_image: $0.00615
fal.aifal_ai/low/2560-x-1440/openai/gpt-image-2.5/sunburst/text-to-image---Image generation; Vision; output_cost_per_image: $0.00615
fal.aifal_ai/low/3840-x-2160/openai/gpt-image-2.5/flare/edit---Image generation; Vision; output_cost_per_image: $0.01113
fal.aifal_ai/low/3840-x-2160/openai/gpt-image-2.5/flare/text-to-image---Image generation; Vision; output_cost_per_image: $0.01113
fal.aifal_ai/low/3840-x-2160/openai/gpt-image-2.5/sunburst/edit---Image generation; Vision; output_cost_per_image: $0.01113
fal.aifal_ai/low/3840-x-2160/openai/gpt-image-2.5/sunburst/text-to-image---Image generation; Vision; output_cost_per_image: $0.01113
fal.aifal_ai/max/1024-x-1024/openai/gpt-image-2.5/flare/edit---Image generation; Vision; output_cost_per_image: $0.21072
fal.aifal_ai/max/1024-x-1024/openai/gpt-image-2.5/flare/text-to-image---Image generation; Vision; output_cost_per_image: $0.21072
fal.aifal_ai/max/1024-x-1024/openai/gpt-image-2.5/sunburst/edit---Image generation; Vision; output_cost_per_image: $0.21072
fal.aifal_ai/max/1024-x-1024/openai/gpt-image-2.5/sunburst/text-to-image---Image generation; Vision; output_cost_per_image: $0.21072
fal.aifal_ai/max/1024-x-1536/openai/gpt-image-2.5/flare/edit---Image generation; Vision; output_cost_per_image: $0.16464
fal.aifal_ai/max/1024-x-1536/openai/gpt-image-2.5/flare/text-to-image---Image generation; Vision; output_cost_per_image: $0.16464
fal.aifal_ai/max/1024-x-1536/openai/gpt-image-2.5/sunburst/edit---Image generation; Vision; output_cost_per_image: $0.16464
fal.aifal_ai/max/1024-x-1536/openai/gpt-image-2.5/sunburst/text-to-image---Image generation; Vision; output_cost_per_image: $0.16464
fal.aifal_ai/max/1024-x-768/openai/gpt-image-2.5/flare/edit---Image generation; Vision; output_cost_per_image: $0.14445
fal.aifal_ai/max/1024-x-768/openai/gpt-image-2.5/flare/text-to-image---Image generation; Vision; output_cost_per_image: $0.14445
fal.aifal_ai/max/1024-x-768/openai/gpt-image-2.5/sunburst/edit---Image generation; Vision; output_cost_per_image: $0.14445
fal.aifal_ai/max/1024-x-768/openai/gpt-image-2.5/sunburst/text-to-image---Image generation; Vision; output_cost_per_image: $0.14445
fal.aifal_ai/max/1920-x-1080/openai/gpt-image-2.5/flare/edit---Image generation; Vision; output_cost_per_image: $0.1584
fal.aifal_ai/max/1920-x-1080/openai/gpt-image-2.5/flare/text-to-image---Image generation; Vision; output_cost_per_image: $0.1584
fal.aifal_ai/max/1920-x-1080/openai/gpt-image-2.5/sunburst/edit---Image generation; Vision; output_cost_per_image: $0.1584
fal.aifal_ai/max/1920-x-1080/openai/gpt-image-2.5/sunburst/text-to-image---Image generation; Vision; output_cost_per_image: $0.1584
fal.aifal_ai/max/2560-x-1440/openai/gpt-image-2.5/flare/edit---Image generation; Vision; output_cost_per_image: $0.2211
fal.aifal_ai/max/2560-x-1440/openai/gpt-image-2.5/flare/text-to-image---Image generation; Vision; output_cost_per_image: $0.2211
fal.aifal_ai/max/2560-x-1440/openai/gpt-image-2.5/sunburst/edit---Image generation; Vision; output_cost_per_image: $0.2211
fal.aifal_ai/max/2560-x-1440/openai/gpt-image-2.5/sunburst/text-to-image---Image generation; Vision; output_cost_per_image: $0.2211
fal.aifal_ai/max/3840-x-2160/openai/gpt-image-2.5/flare/edit---Image generation; Vision; output_cost_per_image: $0.40026
fal.aifal_ai/max/3840-x-2160/openai/gpt-image-2.5/flare/text-to-image---Image generation; Vision; output_cost_per_image: $0.40026
fal.aifal_ai/max/3840-x-2160/openai/gpt-image-2.5/sunburst/edit---Image generation; Vision; output_cost_per_image: $0.40026
fal.aifal_ai/max/3840-x-2160/openai/gpt-image-2.5/sunburst/text-to-image---Image generation; Vision; output_cost_per_image: $0.40026
fal.aifal_ai/medium/1024-x-1024/openai/gpt-image-2.5/flare/edit---Image generation; Vision; output_cost_per_image: $0.01317
fal.aifal_ai/medium/1024-x-1024/openai/gpt-image-2.5/flare/text-to-image---Image generation; Vision; output_cost_per_image: $0.01317
fal.aifal_ai/medium/1024-x-1024/openai/gpt-image-2.5/sunburst/edit---Image generation; Vision; output_cost_per_image: $0.01317
fal.aifal_ai/medium/1024-x-1024/openai/gpt-image-2.5/sunburst/text-to-image---Image generation; Vision; output_cost_per_image: $0.01317
fal.aifal_ai/medium/1024-x-1536/openai/gpt-image-2.5/flare/edit---Image generation; Vision; output_cost_per_image: $0.01029
fal.aifal_ai/medium/1024-x-1536/openai/gpt-image-2.5/flare/text-to-image---Image generation; Vision; output_cost_per_image: $0.01029
fal.aifal_ai/medium/1024-x-1536/openai/gpt-image-2.5/sunburst/edit---Image generation; Vision; output_cost_per_image: $0.01029
fal.aifal_ai/medium/1024-x-1536/openai/gpt-image-2.5/sunburst/text-to-image---Image generation; Vision; output_cost_per_image: $0.01029
fal.aifal_ai/medium/1024-x-768/openai/gpt-image-2.5/flare/edit---Image generation; Vision; output_cost_per_image: $0.00903
fal.aifal_ai/medium/1024-x-768/openai/gpt-image-2.5/flare/text-to-image---Image generation; Vision; output_cost_per_image: $0.00903
fal.aifal_ai/medium/1024-x-768/openai/gpt-image-2.5/sunburst/edit---Image generation; Vision; output_cost_per_image: $0.00903
fal.aifal_ai/medium/1024-x-768/openai/gpt-image-2.5/sunburst/text-to-image---Image generation; Vision; output_cost_per_image: $0.00903
fal.aifal_ai/medium/1920-x-1080/openai/gpt-image-2.5/flare/edit---Image generation; Vision; output_cost_per_image: $0.01029
fal.aifal_ai/medium/1920-x-1080/openai/gpt-image-2.5/flare/text-to-image---Image generation; Vision; output_cost_per_image: $0.01029
fal.aifal_ai/medium/1920-x-1080/openai/gpt-image-2.5/sunburst/edit---Image generation; Vision; output_cost_per_image: $0.01029
fal.aifal_ai/medium/1920-x-1080/openai/gpt-image-2.5/sunburst/text-to-image---Image generation; Vision; output_cost_per_image: $0.01029
fal.aifal_ai/medium/2560-x-1440/openai/gpt-image-2.5/flare/edit---Image generation; Vision; output_cost_per_image: $0.01434
fal.aifal_ai/medium/2560-x-1440/openai/gpt-image-2.5/flare/text-to-image---Image generation; Vision; output_cost_per_image: $0.01434
fal.aifal_ai/medium/2560-x-1440/openai/gpt-image-2.5/sunburst/edit---Image generation; Vision; output_cost_per_image: $0.01434
fal.aifal_ai/medium/2560-x-1440/openai/gpt-image-2.5/sunburst/text-to-image---Image generation; Vision; output_cost_per_image: $0.01434
fal.aifal_ai/medium/3840-x-2160/openai/gpt-image-2.5/flare/edit---Image generation; Vision; output_cost_per_image: $0.02595
fal.aifal_ai/medium/3840-x-2160/openai/gpt-image-2.5/flare/text-to-image---Image generation; Vision; output_cost_per_image: $0.02595
fal.aifal_ai/medium/3840-x-2160/openai/gpt-image-2.5/sunburst/edit---Image generation; Vision; output_cost_per_image: $0.02595
fal.aifal_ai/medium/3840-x-2160/openai/gpt-image-2.5/sunburst/text-to-image---Image generation; Vision; output_cost_per_image: $0.02595
fal.aifal_ai/minimax/h3/reference-to-video---Video generation; output_cost_per_second: $0.13; output_cost_per_second_480p: $0.05; output_cost_per_second_768p: $0.06
fal.aifal_ai/minimax/h3/text-to-video---Video generation; output_cost_per_second: $0.13; output_cost_per_second_480p: $0.05; output_cost_per_second_768p: $0.06
fal.aifal_ai/openai/gpt-image-2.5/flare/edit---Image generation; Vision; output_cost_per_image: $0.03612
fal.aifal_ai/openai/gpt-image-2.5/flare/text-to-image---Image generation; Vision; output_cost_per_image: $0.03612
fal.aifal_ai/openai/gpt-image-2.5/sunburst/edit---Image generation; Vision; output_cost_per_image: $0.03612
fal.aifal_ai/openai/gpt-image-2.5/sunburst/text-to-image---Image generation; Vision; output_cost_per_image: $0.03612
fal.aifal_ai/xhigh/1024-x-1024/openai/gpt-image-2.5/flare/edit---Image generation; Vision; output_cost_per_image: $0.09366
fal.aifal_ai/xhigh/1024-x-1024/openai/gpt-image-2.5/flare/text-to-image---Image generation; Vision; output_cost_per_image: $0.09366
fal.aifal_ai/xhigh/1024-x-1024/openai/gpt-image-2.5/sunburst/edit---Image generation; Vision; output_cost_per_image: $0.09366
fal.aifal_ai/xhigh/1024-x-1024/openai/gpt-image-2.5/sunburst/text-to-image---Image generation; Vision; output_cost_per_image: $0.09366
fal.aifal_ai/xhigh/1024-x-1536/openai/gpt-image-2.5/flare/edit---Image generation; Vision; output_cost_per_image: $0.07377
fal.aifal_ai/xhigh/1024-x-1536/openai/gpt-image-2.5/flare/text-to-image---Image generation; Vision; output_cost_per_image: $0.07377
fal.aifal_ai/xhigh/1024-x-1536/openai/gpt-image-2.5/sunburst/edit---Image generation; Vision; output_cost_per_image: $0.07377
fal.aifal_ai/xhigh/1024-x-1536/openai/gpt-image-2.5/sunburst/text-to-image---Image generation; Vision; output_cost_per_image: $0.07377
fal.aifal_ai/xhigh/1024-x-768/openai/gpt-image-2.5/flare/edit---Image generation; Vision; output_cost_per_image: $0.0642
fal.aifal_ai/xhigh/1024-x-768/openai/gpt-image-2.5/flare/text-to-image---Image generation; Vision; output_cost_per_image: $0.0642
fal.aifal_ai/xhigh/1024-x-768/openai/gpt-image-2.5/sunburst/edit---Image generation; Vision; output_cost_per_image: $0.0642
fal.aifal_ai/xhigh/1024-x-768/openai/gpt-image-2.5/sunburst/text-to-image---Image generation; Vision; output_cost_per_image: $0.0642
fal.aifal_ai/xhigh/1920-x-1080/openai/gpt-image-2.5/flare/edit---Image generation; Vision; output_cost_per_image: $0.07041
fal.aifal_ai/xhigh/1920-x-1080/openai/gpt-image-2.5/flare/text-to-image---Image generation; Vision; output_cost_per_image: $0.07041
fal.aifal_ai/xhigh/1920-x-1080/openai/gpt-image-2.5/sunburst/edit---Image generation; Vision; output_cost_per_image: $0.07041
fal.aifal_ai/xhigh/1920-x-1080/openai/gpt-image-2.5/sunburst/text-to-image---Image generation; Vision; output_cost_per_image: $0.07041
fal.aifal_ai/xhigh/2560-x-1440/openai/gpt-image-2.5/flare/edit---Image generation; Vision; output_cost_per_image: $0.09828
fal.aifal_ai/xhigh/2560-x-1440/openai/gpt-image-2.5/flare/text-to-image---Image generation; Vision; output_cost_per_image: $0.09828
fal.aifal_ai/xhigh/2560-x-1440/openai/gpt-image-2.5/sunburst/edit---Image generation; Vision; output_cost_per_image: $0.09828
fal.aifal_ai/xhigh/2560-x-1440/openai/gpt-image-2.5/sunburst/text-to-image---Image generation; Vision; output_cost_per_image: $0.09828
fal.aifal_ai/xhigh/3840-x-2160/openai/gpt-image-2.5/flare/edit---Image generation; Vision; output_cost_per_image: $0.1779
fal.aifal_ai/xhigh/3840-x-2160/openai/gpt-image-2.5/flare/text-to-image---Image generation; Vision; output_cost_per_image: $0.1779
fal.aifal_ai/xhigh/3840-x-2160/openai/gpt-image-2.5/sunburst/edit---Image generation; Vision; output_cost_per_image: $0.1779
fal.aifal_ai/xhigh/3840-x-2160/openai/gpt-image-2.5/sunburst/text-to-image---Image generation; Vision; output_cost_per_image: $0.1779
xAIxai/grok-4.20-03091,000,000$1.25$2.5Chat; Reasoning; Vision; Tool calling; Prompt caching; Structured output; Web search
xAIxai/grok-4.20-beta1,000,000$1.25$2.5Chat; Reasoning; Vision; Tool calling; Prompt caching; Structured output; Web search
xAIxai/grok-4.20-beta-03091,000,000$1.25$2.5Chat; Reasoning; Vision; Tool calling; Prompt caching; Structured output; Web search
xAIxai/grok-4.20-beta-latest1,000,000$1.25$2.5Chat; Reasoning; Vision; Tool calling; Prompt caching; Structured output; Web search
xAIxai/grok-4.20-beta-latest-non-reasoning1,000,000$1.25$2.5Chat; Vision; Tool calling; Prompt caching; Structured output; Web search
xAIxai/grok-4.20-beta-latest-reasoning1,000,000$1.25$2.5Chat; Reasoning; Vision; Tool calling; Prompt caching; Structured output; Web search
xAIxai/grok-4.20-beta-non-reasoning1,000,000$1.25$2.5Chat; Vision; Tool calling; Prompt caching; Structured output; Web search
xAIxai/grok-4.20-beta-reasoning1,000,000$1.25$2.5Chat; Reasoning; Vision; Tool calling; Prompt caching; Structured output; Web search
xAIxai/grok-4.20-experimental-beta-03041,000,000$1.25$2.5Chat; Reasoning; Vision; Tool calling; Prompt caching; Structured output; Web search
xAIxai/grok-4.20-experimental-beta-0304-non-reasoning1,000,000$1.25$2.5Chat; Vision; Tool calling; Prompt caching; Structured output; Web search
xAIxai/grok-4.20-experimental-beta-0304-reasoning1,000,000$1.25$2.5Chat; Reasoning; Vision; Tool calling; Prompt caching; Structured output; Web search
xAIxai/grok-4.20-experimental-beta-latest1,000,000$1.25$2.5Chat; Reasoning; Vision; Tool calling; Prompt caching; Structured output; Web search
xAIxai/grok-4.20-experimental-beta-non-reasoning-latest1,000,000$1.25$2.5Chat; Vision; Tool calling; Prompt caching; Structured output; Web search
xAIxai/grok-4.20-experimental-beta-reasoning-latest1,000,000$1.25$2.5Chat; Reasoning; Vision; Tool calling; Prompt caching; Structured output; Web search
xAIxai/grok-4.20-multi-agent-beta-latest1,000,000$1.25$2.5Responses; Reasoning; Vision; Prompt caching; Structured output; Web search
xAIxai/grok-4.20-multi-agent-experimental-beta-03041,000,000$1.25$2.5Responses; Reasoning; Vision; Prompt caching; Structured output; Web search
xAIxai/grok-4.20-multi-agent-experimental-beta-latest1,000,000$1.25$2.5Responses; Reasoning; Vision; Prompt caching; Structured output; Web search
xAIxai/grok-4.20-non-reasoning-gv21,000,000$1.25$2.5Chat; Vision; Tool calling; Prompt caching; Structured output; Web search
xAIxai/grok-4.20-reasoning-gv21,000,000$1.25$2.5Chat; Reasoning; Vision; Tool calling; Prompt caching; Structured output; Web search
xAIxai/grok-4.7500,000$2$6Chat; Reasoning; Vision; Tool calling; Prompt caching; Structured output; Web search

Updated pricing (73 models)​

Provider / modelChanged token prices (USD per 1M tokens)
anthropic.claude-mythos-previewInput: $0 to $27.5; Output: $0 to $137.5; Cache read: not set to $2.75; Cache write: not set to $34.375
azure/eu/gpt-5.6Input: $5.5 to $4.4; Output: $33 to $22; Cache read: $0.55 to $0.44; Cache write: $6.875 to $5.5
azure/gpt-5.6Input: $5 to $4; Output: $30 to $20; Cache read: $0.5 to $0.4; Cache write: $6.25 to $5
azure/us/gpt-5.6Input: $5.5 to $4.4; Output: $33 to $22; Cache read: $0.55 to $0.44; Cache write: $6.875 to $5.5
baseten/openai/gpt-oss-120bCache read: not set to $0.1
baseten/zai-org/GLM-4.7Cache read: not set to $0.12
bedrock/eu-west-3/mistral.mistral-large-2402-v1:0Input: $10.4 to $5.2; Output: $31.2 to $15.6
bedrock/us-east-1/mistral.mistral-large-2402-v1:0Input: $8 to $4; Output: $24 to $12
bedrock/us-west-2/mistral.mistral-large-2402-v1:0Input: $8 to $4; Output: $24 to $12
bedrock_mantle/openai.gpt-daybreak-blue-5.6-solInput: $5.5 to $4.4; Output: $33 to $22; Cache read: $0.55 to $0.44; Cache write: $6.875 to $5.5
cohere.command-text-v14Input: $1.5 to $1
deepseek/deepseek-coderCache read: not set to $0.014
deepseek/deepseek-r1Cache read: not set to $0.14
deepseek/deepseek-v3.2Cache read: not set to $0.028
eu.anthropic.claude-opus-4-5-20251101-v1:0Input: $5 to $5.5; Output: $25 to $27.5; Cache read: $0.5 to $0.55; Cache write: $6.25 to $6.875
fireworks_ai/accounts/fireworks/models/minimax-m2p7Input: $0.3 to not set; Output: $1.2 to not set; Cache read: $0.06 to not set
fireworks_ai/accounts/fireworks/routers/kimi-k3-usInput: $3.3 to $4.5; Output: $16.5 to $22.5; Cache read: $0.33 to $0.45
fireworks_ai/kimi-k3-usInput: $3.3 to $4.5; Output: $16.5 to $22.5; Cache read: $0.33 to $0.45
gemini-2.5-flash-imageCache read: $0.03 to not set
gemini-3-pro-image-previewCache read: not set to $0.2
gemini-3.1-flash-image-previewCache read: not set to $0.05
gemini-live-2.5-flash-preview-native-audio-09-2025Cache read: $0.075 to not set
gemini/gemini-live-2.5-flash-preview-native-audio-09-2025Cache read: $0.075 to not set
gemini/gemini-robotics-er-2-streaming-previewInput: $2 to $1; Output: $10 to $5
mistral.mistral-large-2402-v1:0Input: $8 to $4; Output: $24 to $12
openrouter/deepseek/deepseek-r1Cache read: not set to $0.14
openrouter/deepseek/deepseek-v3.2-expCache read: not set to $0.02
openrouter/deepseek/deepseek-v4-flashInput: $0.03724 to $0.049; Output: $0.07448 to $0.098; Cache read: $0.007448 to $0.0098
openrouter/deepseek/deepseek-v4-flash-0731Input: $0.04 to $0.03; Output: $0.08 to $0.32
openrouter/deepseek/deepseek-v4-flash-vision-expInput: $0.2156 to $0.22; Output: $0.6468 to $0.66; Cache read: $0.00686 to $0.007
openrouter/deepseek/deepseek-v4-proInput: $0.422298 to $0.844944; Output: $0.844596 to $1.689888; Cache read: $0.035192 to $0.070412
openrouter/deepseek/deepseek-v4-pro-0813Input: $0.57816 to $0.462; Output: $1.73448 to $1.386; Cache read: $0.018396 to $0.0154
openrouter/meta/muse-glimmer-30bInput: $0.35 to $0.3; Output: $1.5 to $1.2
openrouter/minimax/minimax-m2Input: $0.255 to $0.3; Output: $1.02 to $1.2
openrouter/mistralai/mistral-large-2512Input: $0.55 to $0.5; Output: $1.65 to $1.5; Cache read: $0.055 to $0.05
openrouter/moonshotai/kimi-k2.7-codeInput: $0.7062 to $0.6562; Output: $3.21 to $3.3
openrouter/moonshotai/kimi-k3Input: $1.7 to $3; Output: $8.5 to $15; Cache read: $0.17 to $0.3
openrouter/moonshotai/kimi-k3:batchInput: $3 to $2.28; Output: $15 to $11.4; Cache read: $0.3 to $0.228
openrouter/nvidia/nemotron-3-nano-30b-a3bInput: $0.06 to $0.05; Output: $0.24 to $0.2
openrouter/nvidia/nemotron-3.5-lightningInput: $0.07 to $0.08
openrouter/openai/gpt-oss-120b:batchInput: $0.15 to $0.0296; Output: $0.6 to $0.136
openrouter/openai/gpt-oss-20bInput: $0.03 to $0.018; Output: $0.13 to $0.09
openrouter/qwen/qwen3-30b-a3b-instruct-2507Input: $0.04815 to $0.1; Output: $0.19305 to $0.3
openrouter/qwen/qwen3-next-80b-a3b-instructInput: $0.09 to $0.1
openrouter/qwen/qwen3.6-27bInput: $0.3 to $0.32; Output: $2 to $2.7; Cache read: $0.03 to $0.15
openrouter/qwen/qwen3.6-35b-a3bInput: $0.1 to $0.15; Output: $0.9 to $1
openrouter/z-ai/glm-5.2Input: $0.5544 to $0.6496; Output: $1.7424 to $2.0416; Cache read: $0.10296 to $0.12064
openrouter/z-ai/glm-5.3Input: $0.896 to $1.4; Output: $2.816 to $4.4; Cache read: $0.1664 to $0.26
openrouter/z-ai/glm-5.3-flashInput: $0.09 to $0.045; Output: $0.3 to $0.6; Cache read: $0.018 to $0.0285
openrouter/z-ai/glm-5.3-flash:batchInput: $0.075 to $0.06; Output: $0.25 to $0.2; Cache read: $0.015 to $0.012
openrouter/z-ai/glm-5.3-flashxCache read: $0.075 to $0.09
openrouter/z-ai/glm-5.3:batchInput: $0.7 to $0.45; Output: $2.2 to $2; Cache read: $0.13 to $0.1
openrouter/~anthropic/claude-opus-latestInput: $5 to $4; Output: $25 to $20; Cache read: $0.5 to $0.2; Cache write: $6.25 to $5
openrouter/~deepseek/deepseek-flash-latestInput: $0.13 to $0.3; Output: $0.52 to $1.2; Cache read: $0.0026 to $0.006
openrouter/~deepseek/deepseek-pro-latestInput: $0.57816 to $1.32; Output: $1.73448 to $3.96; Cache read: $0.018396 to $0.044
openrouter/~deepseek/deepseek-v4-flash-latestOutput: $0.08 to $0.64
openrouter/~moonshotai/kimi-latestInput: $1.7 to $3; Output: $8.5 to $15; Cache read: $0.17 to $0.3
openrouter/~openai/gpt-luna-latestInput: $0.2 to $0.1; Output: $1.2 to $0.5; Cache read: $0.02 to $0.01; Cache write: $0.25 to $0.125
openrouter/~x-ai/grok-latestInput: $2 to $1.6; Output: $6 to $4.8; Cache read: $0.5 to $0.4
openrouter/~z-ai/glm-flash-latestInput: $0.075 to $0.15; Output: $0.25 to $0.5; Cache read: $0.015 to $0.05
openrouter/~z-ai/glm-latestInput: $0.8442 to $0.6538; Output: $2.6532 to $2.0548; Cache read: $0.15678 to $0.12142
together_ai/Qwen/Qwen3.7-MaxInput: $2.5 to $1.5; Output: $7.5 to $4.5; Cache read: $0.5 to $0.3
together_ai/Qwen/Qwen3.8-FlashInput: $0.15 to $0.09; Output: $0.47 to $0.282
vertex_ai/deepseek-ai/deepseek-v3.1-maasCache read: not set to $0.06
vertex_ai/deepseek-ai/deepseek-v3.2-maasCache read: not set to $0.056
vertex_ai/gemini-2.5-flash-imageCache read: $0.03 to not set
vertex_ai/gemini-3-pro-image-previewCache read: not set to $0.2
vertex_ai/gemini-3.1-flash-image-previewCache read: not set to $0.05
vertex_ai/google/gemma-4-26b-a4b-it-maasCache read: not set to $0.015
vertex_ai/minimaxai/minimax-m2-maasCache read: not set to $0.03
vertex_ai/moonshotai/kimi-k2-thinking-maasCache read: not set to $0.06
vertex_ai/qwen/qwen3-coder-480b-a35b-instruct-maasCache read: not set to $0.022
vertex_ai/zai-org/glm-4.7-maasCache read: not set to $0.06

The registry also updates capability flags, context/output limits, non-token rates, and deprecation dates on 630 more entries

Removed catalog entries (275)

1024-x-1024/dall-e-2, 256-x-256/dall-e-2, 512-x-512/dall-e-2, amazon.nova-sonic-v1:0, anthropic.claude-3-haiku-20240307-v1:0, anthropic.claude-3-sonnet-20240229-v1:0, apac.anthropic.claude-3-5-sonnet-20240620-v1:0, apac.anthropic.claude-3-5-sonnet-20241022-v2:0, apac.anthropic.claude-3-haiku-20240307-v1:0, apac.anthropic.claude-3-sonnet-20240229-v1:0, azure/eu/o1-preview-2024-09-12, azure/global/gpt-5.1-chat, azure/gpt-3.5-turbo-0125, azure/gpt-35-turbo-0125, azure/gpt-35-turbo-1106, azure/gpt-5-chat-latest, azure/gpt-5.1-chat-2025-11-13, azure/gpt-5.2-chat-2025-12-11, azure/o1-preview, azure/o1-preview-2024-09-12, azure/us/o1-preview-2024-09-12, azure_ai/Llama-3.2-11B-Vision-Instruct, azure_ai/Llama-3.2-90B-Vision-Instruct, azure_ai/MAI-Image-2e, azure_ai/Meta-Llama-3.1-405B-Instruct, azure_ai/Meta-Llama-3.1-8B-Instruct, azure_ai/claude-opus-4-1, azure_ai/cohere-rerank-v3.5, azure_ai/global/grok-3, azure_ai/global/grok-3-mini, azure_ai/mistral-document-ai-2505, bedrock/us-gov-east-1/anthropic.claude-3-5-sonnet-20240620-v1:0, bedrock/us-gov-east-1/anthropic.claude-3-haiku-20240307-v1:0, bedrock/us-gov-west-1/anthropic.claude-3-5-sonnet-20240620-v1:0, bedrock/us-gov-west-1/anthropic.claude-3-7-sonnet-20250219-v1:0, bedrock/us-gov-west-1/anthropic.claude-3-haiku-20240307-v1:0, cerebras/zai-glm-4.6, cerebras/zai-glm-4.7, chatgpt-4o-latest, claude-3-7-sonnet-20250219, claude-3-haiku-20240307, claude-3-opus-20240229, claude-4-opus-20250514, claude-4-sonnet-20250514, claude-opus-4-1, claude-opus-4-1-20250805, claude-opus-4-20250514, claude-sonnet-4-20250514, codex-mini-latest, cohere.command-r-plus-v1:0, cohere.command-r-v1:0, command, command-light, command-r, command-r-plus, dall-e-2, dall-e-3, databricks/databricks-claude-3-7-sonnet, databricks/databricks-gpt-5-1-codex-max, databricks/databricks-gpt-5-1-codex-mini, databricks/databricks-gpt-5-2-codex, databricks/databricks-llama-2-70b-chat, databricks/databricks-meta-llama-3-1-405b-instruct, databricks/databricks-meta-llama-3-70b-instruct, databricks/databricks-mixtral-8x7b-instruct, databricks/databricks-mpt-30b-instruct, databricks/databricks-mpt-7b-instruct, deepinfra/google/gemini-2.0-flash-001, embed-english-light-v2.0, embed-english-v2.0, embed-multilingual-v2.0, eu.anthropic.claude-3-haiku-20240307-v1:0, eu.anthropic.claude-3-sonnet-20240229-v1:0, fireworks_ai/deepseek-v4-pro, fireworks_ai/minimax-m2p7, friendliai/LGAI-EXAONE/K-EXAONE-2.0-750B-A37B, gemini-2.0-flash, gemini-2.0-flash-001, gemini-2.0-flash-lite, gemini-2.0-flash-lite-001, gemini-2.5-flash-lite-preview-06-17, gemini-3-pro-preview, gemini/gemini-1.5-flash, gemini/gemini-2.0-flash, gemini/gemini-2.0-flash-001, gemini/gemini-2.0-flash-lite, gemini/gemini-2.0-flash-lite-001, gemini/gemini-2.5-flash-lite-preview-06-17, gemini/gemini-2.5-flash-lite-preview-09-2025, gemini/gemini-2.5-flash-preview-09-2025, gemini/gemini-3-pro-preview, gemini/gemini-robotics-er-1.5-preview, gemini/gemini-robotics-er-1.6-preview, gemini/imagen-3.0-generate-002, gemini/imagen-4.0-fast-generate-001, gemini/imagen-4.0-generate-001, gemini/imagen-4.0-ultra-generate-001, gemini/veo-2.0-generate-001, gpt-4-0125-preview, gpt-4-0314, gpt-4-turbo-preview, gpt-4o-audio-preview, gpt-4o-mini-audio-preview, gpt-4o-mini-realtime-preview, gpt-4o-mini-search-preview-2025-03-11, gpt-4o-mini-tts-2025-03-20, gpt-4o-realtime-preview, gpt-4o-realtime-preview-2024-12-17, gpt-4o-realtime-preview-2025-06-03, gpt-4o-search-preview-2025-03-11, gpt-audio-mini-2025-10-06, gpt-realtime-mini-2025-10-06, groq/gemma-7b-it, groq/llama-3.1-8b-instant, groq/llama-3.3-70b-versatile, groq/meta-llama/llama-4-maverick-17b-128e-instruct, groq/meta-llama/llama-4-scout-17b-16e-instruct, groq/meta-llama/llama-guard-4-12b, groq/moonshotai/kimi-k2-instruct-0905, groq/playai-tts, groq/qwen/qwen3-32b, groq/qwen/qwen3.6-27b, hd/1024-x-1024/dall-e-3, hd/1024-x-1792/dall-e-3, hd/1792-x-1024/dall-e-3, mistral/codestral-2405, mistral/devstral-2512, mistral/devstral-medium-2507, mistral/devstral-small-2505, mistral/devstral-small-2507, mistral/labs-devstral-small-2512, mistral/magistral-medium-1-2-2509, mistral/magistral-medium-2506, mistral/magistral-medium-2509, mistral/magistral-small-1-2-2509, mistral/magistral-small-2506, mistral/mistral-large-2402, mistral/mistral-large-2407, mistral/mistral-large-2411, mistral/mistral-medium-2312, mistral/mistral-medium-2505, mistral/mistral-medium-2508, mistral/mistral-medium-3-1-2508, mistral/mistral-ocr-2505-completion, mistral/mistral-small-3-2-2506, mistral/open-codestral-mamba, mistral/open-mistral-7b, mistral/open-mistral-nemo-2407, mistral/open-mixtral-8x22b, mistral/open-mixtral-8x7b, mistral/pixtral-12b-2409, mistral/pixtral-large-2411, moonshot/kimi-k2-0711-preview, moonshot/kimi-k2-0905-preview, moonshot/kimi-k2-thinking, moonshot/kimi-k2-thinking-turbo, moonshot/kimi-k2-turbo-preview, moonshot/kimi-latest, moonshot/kimi-latest-128k, moonshot/kimi-latest-32k, moonshot/kimi-latest-8k, moonshot/kimi-thinking-preview, moonshot/moonshot-v1-128k-0430, moonshot/moonshot-v1-32k-0430, moonshot/moonshot-v1-8k-0430, o3-deep-research-2025-06-26, o4-mini-deep-research-2025-06-26, openrouter/anthropic/claude-opus-4, openrouter/deepseek/deepseek-v4-flash-0731:batch, openrouter/deepseek/deepseek-v4-flash-0731:free, openrouter/deepseek/deepseek-v4-flash-vision-exp:batch, openrouter/deepseek/deepseek-v4-pro-0813:batch, openrouter/google/gemini-2.0-flash-001, openrouter/kwaipilot/kat-coder-pro-v2, openrouter/meta/muse-glimmer-30b:batch, openrouter/minimax/minimax-m3:batch, openrouter/qwen/qwen3.5-9b:batch, openrouter/qwen/qwen3.8-2.4t-a95b:batch, openrouter/stealth/union-alpha, openrouter/thinkingmachines/inkling:batch, openrouter/z-ai/glm-5.2:batch, rerank-english-v2.0, rerank-multilingual-v2.0, sambanova/DeepSeek-R1-Distill-Llama-70B, sambanova/DeepSeek-V3-0324, sambanova/Llama-4-Scout-17B-16E-Instruct, sambanova/Meta-Llama-3.1-405B-Instruct, sambanova/Meta-Llama-3.1-8B-Instruct, sambanova/Meta-Llama-3.2-1B-Instruct, sambanova/Meta-Llama-3.2-3B-Instruct, sambanova/Meta-Llama-Guard-3-8B, sambanova/QwQ-32B, sambanova/Qwen2-Audio-7B-Instruct, sambanova/Qwen3-32B, scaleway/google/gemma-3-27b-it, scaleway/hcompany/holo2-30b-a3b, scaleway/mistralai/devstral-2-123b-instruct-2512, scaleway/mistralai/voxtral-small-24b-2507, standard/1024-x-1024/dall-e-3, standard/1024-x-1792/dall-e-3, standard/1792-x-1024/dall-e-3, text-moderation-007, text-moderation-latest, text-moderation-stable, together_ai/Qwen/Qwen3-235B-A22B-Instruct-2507-tput, together_ai/Qwen/Qwen3-235B-A22B-Thinking-2507, together_ai/Qwen/Qwen3-235B-A22B-fp8-tput, together_ai/deepseek-ai/DeepSeek-R1, together_ai/deepseek-ai/DeepSeek-R1-0528-tput, together_ai/deepseek-ai/DeepSeek-V4-Pro, together_ai/google/gemma-3n-E4B-it, together_ai/intfloat/multilingual-e5-large-instruct, together_ai/meta-llama/Llama-3.2-3B-Instruct-Turbo, together_ai/meta-llama/Llama-3.3-70B-Instruct-Turbo-Free, together_ai/meta-llama/Llama-4-Maverick-17B-128E-Instruct-FP8, together_ai/meta-llama/Llama-Guard-4-12B, together_ai/meta-llama/Meta-Llama-3.1-405B-Instruct-Turbo, together_ai/moonshotai/Kimi-K2-Instruct-0905, together_ai/moonshotai/Kimi-K2.5, together_ai/pearl-ai/gemma-4-31b-it, together_ai/thinkingmachines/Inkling-Small, us-gov.anthropic.claude-3-haiku-20240307-v1:0, us.amazon.nova-premier-v1:0, us.anthropic.claude-3-haiku-20240307-v1:0, us.anthropic.claude-3-sonnet-20240229-v1:0, vercel_ai_gateway/google/gemini-2.0-flash, vercel_ai_gateway/google/gemini-2.0-flash-lite, vertex_ai/claude-3-7-sonnet@20250219, vertex_ai/claude-opus-4, vertex_ai/claude-opus-4-1, vertex_ai/claude-opus-4-1@20250805, vertex_ai/claude-opus-4@20250514, vertex_ai/claude-sonnet-4, vertex_ai/claude-sonnet-4@20250514, vertex_ai/gemini-3.1-flash-live-preview, vertex_ai/gemini-robotics-er-2, vertex_ai/imagegeneration@006, vertex_ai/imagen-3.0-capability-001, vertex_ai/imagen-3.0-fast-generate-001, vertex_ai/imagen-3.0-generate-001, vertex_ai/imagen-3.0-generate-002, vertex_ai/imagen-4.0-fast-generate-001, vertex_ai/imagen-4.0-generate-001, vertex_ai/imagen-4.0-ultra-generate-001, wandb/MiniMaxAI/MiniMax-M2.5, wandb/Qwen/Qwen3-235B-A22B-Instruct-2507, wandb/Qwen/Qwen3-235B-A22B-Thinking-2507, wandb/Qwen/Qwen3-Coder-480B-A35B-Instruct, wandb/deepseek-ai/DeepSeek-R1-0528, wandb/deepseek-ai/DeepSeek-V3-0324, wandb/meta-llama/Llama-4-Scout-17B-16E-Instruct, wandb/microsoft/Phi-4-mini-instruct, wandb/moonshotai/Kimi-K2-Instruct, wandb/zai-org/GLM-4.5, xai/grok-3, xai/grok-3-beta, xai/grok-3-fast-beta, xai/grok-3-fast-latest, xai/grok-3-latest, xai/grok-3-mini, xai/grok-3-mini-beta, xai/grok-3-mini-fast, xai/grok-3-mini-fast-beta, xai/grok-3-mini-fast-latest, xai/grok-3-mini-latest, xai/grok-4, xai/grok-4-0709, xai/grok-4-1-fast, xai/grok-4-1-fast-non-reasoning, xai/grok-4-1-fast-non-reasoning-latest, xai/grok-4-1-fast-reasoning, xai/grok-4-1-fast-reasoning-latest, xai/grok-4-fast-non-reasoning, xai/grok-4-fast-reasoning, xai/grok-4-latest

New providers​

  • Add Nadir intelligent-router provider (nadir/auto) - PR #33227
  • Add Eden AI provider across chat, Responses, Messages, embeddings, audio, images and video - PR #41101

Amazon Bedrock​

  • Strip body params the AWS endpoint rejects - PR #31203
  • Drop unsupported sampling params on converse reasoning models - PR #39834
  • Send s3BucketOwner on batch input and output data config - PR #42262
  • Add us.moonshotai.kimi-k3 pricing and fill the global Kimi K3 entry - PR #42271
  • Forward anthropic-beta headers verbatim on the Claude platform messages path - PR #42275
  • Keep batch S3 credentials out of chat requests and debug logs - PR #42312
  • Sign batch S3 requests with s3_access_key_id and s3_secret_access_key - PR #42342
  • Add bare moonshotai.kimi-k3 cost map entry - PR #42363
  • Send every Mantle beta in the anthropic-beta header on the bedrock/mantle route - PR #42376
  • Price bedrock/mantle/<model> deployments from the base model row - PR #42402
  • Treat blank AWS_S3_* env vars as unset for batch jobs - PR #42528
  • Add Claude Opus 5.5 pricing and capabilities - PR #42588
  • Send json_schema as a forced tool on Claude Opus 4.7 and 4.8 Converse - PR #42644
  • Honour stream_chunk_size in Invoke streaming - PR #42686
  • Route unmapped openai family model ids to converse - PR #42713
  • Add gpt-6-sol and gpt-6-luna model pricing - PR #42746
  • Serve the OpenAI models on bedrock-runtime's native Responses API (internal copy of #38489) - PR #42767
  • Add bare openai.gpt-6-sol and openai.gpt-6-luna cost map rows - PR #42798
  • Add 17 aws-bedrock cost map rows from provider sync - PR #42852
  • Add gpt-5.4 and gpt-5.5 us and global inference profile pricing - PR #42941
  • Extrapolate global cris pricing for gpt-5.4 and gpt-5.5 - PR #42971
  • Map Anthropic batch row params the way real time does - PR #43087

Anthropic​

  • Forward safeguards and anthropic-beta unchanged on native /v1/messages - PR #42152
  • Forward Claude Code safeguards and dangerous-tool-use beta to Bedrock Invoke and Vertex on /v1/messages - PR #42288
  • Keep reasoning_effort a string for targets that stay on chat completions - PR #42401
  • Return 400 instead of 500 when a content list holds a bare string - PR #42420
  • Add Claude Opus 5.5 - PR #42489
  • Keep the replayed prefix byte-stable for preserved thinking on chat completions - PR #42630
  • Preserve MCP tool results in the non-Anthropic Messages bridge - PR #42783

Azure​

  • Bridge gpt-5.4+ function tools with reasoning to the Foundry Responses API - PR #42041
  • Propagate asyncio.CancelledError instead of raising a 500 - PR #42295
  • Add gpt-audio and gpt-realtime alias rows from the Azure model list - PR #42981
  • Update gpt-audio-mini and gpt-5-chat deprecation dates from the retirement schedule - PR #43017

Bedrock Mantle​

  • Serve /v1/messages for Claude models on Mantle's native Anthropic Messages API - PR #42049

fal.ai​

  • Add Seedance 2.5 / 2.0 video generation via fal queue API - PR #41980
  • Add gpt-image-2.5 flare/sunburst, flux/dev and image edits - PR #42095
  • Price images from the dimensions fal returns - PR #42282
  • Add MiniMax H3 text-to-video and reference-to-video - PR #42286
  • Surface fal errors in video status and content instead of completed and generic 500 - PR #42306
  • Add flux-lora-depth image edits and moondream3 chat completions - PR #42334
  • Price non-canonical image sizes from the nearest row and honour dump options - PR #42336
  • Add queue-only /fal_ai pass-through route with spend tracking - PR #42360
  • Handle seconds=auto and oversized sizes for minimax h3 videos - PR #42504
  • Align /fal_ai queue gate with the pricer and normalise resolution type - PR #42505
  • Reuse the status client and headers on the video result probe - PR #42511
  • Honour global api_base for image generation and reject non-string reasoning_effort with 400 - PR #42512
  • Price nano-banana-2 and nano-banana-pro image generations by resolution - PR #43101

Fireworks AI​

  • Route firerouter short names and bill pass-through legs at the routed model's rates - PR #42814

Gemini and Vertex AI​

  • Forward response schema and tool parameters through the generateContent adapter - PR #42067
  • Simplify model version check - PR #42465
  • Add gemini-3.8-flash-tts and gemini-3.8-flash-lite-tts prices - PR #42752
  • Native batch JSONL passthrough with cost tracking - PR #42810
  • Add deprecation dates for retired claude 3 and jamba 1.5 partner models - PR #42867
  • Add gemini-3.8-flash-cyber pricing - PR #42879
  • Keep batch output_file_id null until Vertex reports outputInfo - PR #43030
  • Translate /v1/responses batch rows through the Responses-to-Chat bridge - PR #43042
  • Surface the Gemma container's own error inside a 200 :predict response - PR #43075
  • Stop advertising OpenAI platform-only params on Gemma and Llama routes - PR #43079
  • Return chunk content, extractive text, and structData from search_api vector store hits - PR #43100

GitHub Copilot and ChatGPT​

  • Answer get_api_base for github_copilot and chatgpt without running the login flow - PR #42602

Ollama​

  • Send PNG and JPEG images without requiring Pillow - PR #41979
  • Read the JSON thinking field on non-streaming completions - PR #42838

OpenAI​

  • Forward non-enum reasoning_effort through the Responses bridge instead of dropping it - PR #42452
  • Add GPT-6 Sol and GPT-6 Luna - PR #42515

OpenRouter​

  • Price typesafe/jev-1.13 and add an openrouter decisions pass-through - PR #42301
  • Remove the retired stealth/union-alpha model from the cost map - PR #42386

Qianwen AI Platform​

  • Rename the mainland China brand to Qianwen AI Platform - PR #42284

Together AI​

  • Backfill deprecation_date from Together deprecation history - PR #43135

xAI​

  • Add grok-4.7 to the cost map - PR #42264
  • Accept max_completion_tokens as a supported param - PR #42353

Xiaomi MiMo​

  • Add mimo-v2.6-pro and mimo-v2.6-flash cost map rows with live e2e coverage - PR #42362

Model catalog and pricing​

  • New and updated catalog rows
    • Add MAI-Image-2.5-Pro pricing, fix Fireworks/Together entries, absorb verified open registry PRs, add Groq deprecation and Bedrock regional Qwen3 Next pricing - PR #34941
    • Add xai grok-4.20 aliases and image token prices from /v1/language-models - PR #42384
    • Add Claude Opus 5.5 for Vertex AI and Azure AI - PR #42599
    • Add openrouter/aion-labs/aion-3.5-mini pricing - PR #42743
    • Add Azure Foundry pricing for gpt-6-sol and gpt-6-luna - PR #42747
    • Add openrouter/stealth/space-bunny-alpha - PR #42759
    • Add baseten/zai-org/GLM-5.3-Fast pricing - PR #42764
    • Add together_ai/together/Tev1-4B-experimental - PR #42807
    • Add gemini preview aliases and deep research 04-2026 rows - PR #42833
    • Add openai chat-latest, codex and deep-research rows from the model docs - PR #42834
    • Add vertex ai llama 3.3 70b, veo 2/3, virtual try-on and 2.5 tts rows - PR #42837
    • Add openrouter/openai/gpt-oss-120b:batch from the OpenRouter models API - PR #42847
    • Add gemini lyria-realtime-exp row inherited from lyria-3.5 - PR #42848
    • Add 39 together_ai chat rows priced by the Together models API - PR #42851
    • Add retired azure gpt-5 chat and o1-preview data zone rows - PR #42878
    • Add vertex_ai/gemini-3.8-live - PR #42891
    • Add wandb DeepSeek-V4.1-Flash and gemma-4-26B-A4B-it - PR #42924
    • Add azure realtime, audio and partner model rows - PR #42954
    • Sync azure models, add MAI-Image-2.6, deepseek-v4.1-flash, muse-spark-1.3 - PR #42970
    • Add supports_reasoning to azure/eu/gpt-6-astra - PR #42989
    • Add fireworks glm-5p3 US-only rows and kimi-k3-us priority prices - PR #43092
    • Add fireworks deepseek-v4p1-flash US-only rows - PR #43097
  • Price corrections
    • Registry audit 2026-09-22, absorb open pricing PRs - PR #42543
    • Drop the unpublished cached rate from the Gemini Live preview entries - PR #42651
    • Align bedrock_mantle/openai.gpt-daybreak-blue-5.6-sol with its Bedrock model card - PR #42672
    • Align regional Bedrock Mistral Large 24.02 keys with the AWS pricing page - PR #42684
    • Registry audit 2026-09-23, in-region Bedrock Claude prices - PR #42779
    • Sync openrouter prices from the models API - PR #42832
    • Bedrock bare Claude ids priced at the Global SKU (aws-bedrock sync) - PR #42875
    • Correct gemini robotics er 2 preview audio input price - PR #42877
    • Halve openrouter deepseek-v4-flash-0731 output price - PR #42881
    • Correct fireworks kimi k3 us pricing to the published rate - PR #42884
    • Sync openrouter prices and add fireworks ember-1 - PR #42889
    • Source and chat completions endpoint for bedrock mantle gpt-5.4 and gpt-5.5 - PR #42890
    • Source for bedrock mantle gpt-5.6 luna, sol, terra and grok-4.6 - PR #42898
    • Update openrouter kimi-k2.7-code input price - PR #42932
    • Sync vertex-ai rows (gemma 4 maas cache price, chirp_2) - PR #42942
    • Retirement dates, chatgpt reasoning flags, bing pricing, bedrock mantle and mythos, azure gpt-5.6 alias, anthropic batch rates, new nebius, openrouter and xai rows - PR #42951
    • Sync openrouter prices for deepseek v4 and glm-5.3 - PR #42952
    • Add batch prices for vertex gemini-3.8-flash-cyber - PR #42953
    • Sync openrouter deepseek-v4-pro prices - PR #42956
    • Sync openrouter deepseek-v4-pro prices - PR #42964
    • Sync openrouter deepseek-v4-pro prices - PR #42969
    • Sync openrouter deepseek v4 flash, v4 pro and v4.1 flash prices - PR #42974
    • Sync openrouter deepseek-v4-pro prices - PR #42978
    • Add vertex ai priority audio input prices for gemini flash rows - PR #42980
    • Align openrouter deepseek-v4-pro cache hit price with cache read - PR #42985
    • Drop stale cache_hit field from openrouter deepseek-v4-pro - PR #42994
    • Add vertex cache read, batch and above 200k prices for gemini image preview rows - PR #42995
    • Add video and reasoning output prices to vertex gemini-omni-1.1-flash - PR #43036
    • Drop the priority input price from vertex gemini-2.5-flash-image - PR #43050
    • Add vertex priority prices for gemini-3-pro-image-preview and batch price for gemini-embedding-001 - PR #43069
    • Add video input price to gemini 3.8 live rows - PR #43073
    • Drop cache read price from vertex gemini-2.5-flash-image - PR #43078
    • Sync openrouter deepseek v4 and glm-5.3-flash prices, add mistral-large-2512 - PR #43090
    • Sync gemini priority, flex and video token prices from the Gemini API pricing page - PR #43091
    • Add gemini tts batch output prices from the Gemini API pricing page - PR #43103
    • Add openai cached image input prices from the pricing page - PR #43143
  • Deprecation and retirement dates
    • Add azure_ai gpt-image-2 and groq llama-guard-3-8b deprecation dates - PR #42738
    • Add deprecation_date to claude-mythos-preview from the Anthropic deprecations page - PR #42845
    • Add the sora-2-pro shutdown date to the sora-2-pro-high-res rows - PR #42846
    • Add fireworks deprecation dates for kimi k2.6 fast, kimi k2.7 code fast and glm 5.2 fast us - PR #42849
    • Add the June 1, 2026 retirement date to the vertex_ai gemini-2.0-flash rows - PR #42850
    • Add fireworks deprecation date for glm 5.2 fast serverless rows - PR #42866
    • Add openai deprecation date for gpt-5-chat-latest and gpt-5-chat - PR #42872
    • Add azure gpt-4o-mini-transcribe and gpt-4o-mini-tts deprecation dates - PR #42873
    • Add fireworks 2026-09-25 deprecation dates for glm 5.2, kimi k2.6, kimi k2.7 code, deepseek v4 and muse glimmer rows - PR #42874
    • Sync vertex-ai deprecation dates from Vertex model lifecycle pages - PR #42882
    • Add azure gpt-realtime-mini deprecation date - PR #42883
    • Add azure gpt-realtime-mini-2025-10-06 deprecation date - PR #42885
    • Add azure deprecation dates for gpt-6 and gpt-realtime-whisper - PR #42897
    • Add azure deprecation dates for regional gpt-6 rows - PR #42933
    • Update azure gpt-4.1-nano and gpt-4o-2024-05-13 retirement dates - PR #42947
    • Take the later azure deprecation date for gpt-4.1-nano, gpt-4o-transcribe and gpt-realtime-2.1 - PR #42986
    • Add the Vertex shutdown date to gemini-2.5-flash-native-audio - PR #43024
    • Add azure deprecation dates from the Models API for five realtime and transcribe rows - PR #43037
    • Move azure gpt-realtime-2.1-mini deprecation date to the later Models API date - PR #43058
    • Add openai deprecation dates from the deprecations page - PR #43102
    • Add azure retirement dates from the retired Foundry models page - PR #43104
    • Add computer-use-preview deprecation date from the openai deprecations page - PR #43116
    • Add azure retirement dates for command-r-plus and gpt-4 - PR #43117
    • Add together-ai deprecation dates for gpt-oss-20b and gemma-4-31B-it - PR #43127
  • Removals and cleanup
    • Remove models past their deprecation date - PR #42435
    • Remove retired models flagged by the provider sync - PR #42521
    • Drop duplicate cache_read_input_token_cost_batches keys from the price maps - PR #42623
  • Automated provider price syncs (78 PRs: OpenRouter 51, AWS Bedrock 11, OpenAI 3, Baseten 3, Vertex AI 3, Azure 2, Google Gemini 2, Together AI 1, Fireworks AI 1, xAI 1)

LLM API Endpoints​

Responses API​

  • Stream one lifecycle across MCP auto-execute rounds - PR #40121
  • Stop agentic follow-up from passing request params twice - PR #41560
  • Patch custom_tool_call_output in place on guardrail write-back - PR #41561
  • Forward safety_identifier through the chat completion bridge - PR #42193
  • Build follow-up kwargs in one place so no executor can repeat a request param - PR #42307
  • Drop client_metadata and merge system messages for Databricks chat-only models - PR #42390

Agents and Agent-to-Agent​

  • Attach access groups to agents and enforce them for models, MCP servers and agent calls - PR #41634
  • Send message/stream for Bedrock AgentCore streaming requests - PR #42239
  • Add optional per-agent kill switch webhook - PR #42841

OCR​

  • Remove the Python OCR execution path and require the Rust route - PR #43081

Realtime and Audio​

  • Surface an upstream handshake refusal as an error event and policy close - PR #42388
  • Keep config-defined vector stores listed and read-only - PR #42574

Pass-through endpoints​

  • Add TinyFish Agent API passthrough with per-step billing - PR #41099
  • Log upstream 4xx/5xx error bodies and carry them into the failure hook - PR #42695

General​

  • Stop /utils/transform_request from calling the provider and blocking the event loop - PR #33954
  • Prefilled GitHub issue link on unmapped internal errors - PR #42065
  • Keep an explicit provider prompt_tokens=0 or completion_tokens=0 in streamed usage - PR #42323
  • Let a later usage event zero out stale cache counts (#40736) - PR #42330
  • Opt-in litellm_call_id in JSON error bodies - PR #42391
  • Add stream and safe config flags to the bug report link - PR #42428
  • Stop a nested additional_drop_params entry from crashing openai-compatible calls - PR #42492
  • Point the blocked-address remediation at litellm_settings - PR #42508
  • Isolate callback errors in async_post_call_success_deployment_hook - PR #42535
  • Repair seven regressions caught by CircleCI on main - PR #42640
  • Stop stream_chunk_size reaching provider request bodies - PR #42664

Management Endpoints / UI​

Admin UI​

  • Surface the owner's user budget on keys without their own budget - PR #38220
  • Add upgrade banner with latest release changelog stats - PR #40429
  • Let the Create Key user picker find users by user_id, not just email - PR #41687
  • Add internal-user savings and auto-router usage - PR #42026
  • Show prompt caching requests and net savings - PR #42055
  • Show Capability and FUSE v2 routing forecasts - PR #42057
  • Expose remaining complexity router advanced settings - PR #42293
  • Add per-user breakdown to team usage export - PR #42367
  • Rename reminder markers to Ignore Custom Tags - PR #42370
  • Add native LiteAdmin assistant - PR #42443
  • Add span type filter to request logs - PR #42491
  • Simplify auto-router setup and clarify feature limits - PR #42625
  • Show ten prompt caching requests per page - PR #42638
  • Prefer native providers in auto-router presets - PR #42639
  • Keep per-user MCP credentials updatable and clearable after setup - PR #42652
  • Keep untouched stored auto-router booleans and reminder marker casing on save - PR #42703
  • Hide LiteAdmin in Playground and add admin preference - PR #42755
  • Make the audit log detail drawer wider and resizable - PR #42808
  • Offer reset of custom member budgets when team default changes - PR #42835
  • Configure prompt caching request rows per page - PR #42842
  • Explain unbackfilled key lifetime spend and ship a backfill script - PR #42967
  • Stop following streamed tokens, add jump to bottom button - PR #42968
  • Pass is_proxy_admin for proxy admins on the models page team drill-in - PR #43003
  • Group cost optimization cache leakage by model group - PR #43008

Keys, teams and authentication​

  • Invalidate cached object permissions on key update - PR #36719
  • Accept token_id as an alternative to the plaintext key - PR #39578
  • Count team unified access group MCP servers when validating key MCP grants - PR #41231
  • Use the user's own budget as the ceiling for UI session personal keys - PR #41588
  • Show whether a member follows the team default budget and allow resetting to it - PR #41906
  • Refuse to start with an unset, empty, or publicly known master key - PR #42019
  • Fail closed when the team membership lookup hits a db outage - PR #42036
  • Reject deactivated JWT users and refresh cached status - PR #42064
  • Breached password detection, self-service change-password and forced password reset - PR #42278
  • Evict the cached user row when SCIM or /user/delete removes a user - PR #42315
  • Fail closed when the JWT single-team fallback or compact editor membership read hits a DB outage - PR #42344
  • Let jwt team_allowed_routes paths grant auth=true passthrough - PR #42346
  • Surface a database outage from the user read as 503 no_db_connection - PR #42399
  • Answer 503 no_db_connection on management routes when the caller's user read hits a database outage - PR #42410
  • Enforce disable_custom_api_keys from general_settings - PR #42437
  • Write key deleted audit logs for cascade and alias key deletions - PR #42446
  • Revoke UI session tokens on logout and password change - PR #42463
  • Configurable key_alias_pattern for key generate, update, and regenerate - PR #42553
  • Let team admins update member key budgets when enabled - PR #42555
  • Authorize key model aliases the same way as team aliases - PR #43049

Proxy configuration​

  • Keep config-defined deployments when a config read returns no model_list - PR #41505
  • Report the source of alerting, UI and router settings on read - PR #41788
  • Drop cost-map metadata echoed back on model save - PR #41944
  • Remove the dead telemetry flag from the SDK, proxy CLI and configs - PR #42071
  • Add GET /utils/model_info to look up cost map info for unregistered models - PR #42121
  • Detach stored credential when model editor selects None - PR #42291
  • Admin-only /debug/report sharing the bug report environment - PR #42440
  • Never render credential-bearing config keys in the bug report - PR #42493
  • Make the lazy OpenAPI snapshot byte-identical on every Python version - PR #42519
  • Validate model credential name only when it changes - PR #42701
  • Document request body and response schemas for the Responses API in OpenAPI - PR #42802
  • Honor model_info.discoverable on the model listing endpoints - PR #42825
  • List key and team model aliases in GET /v1/models - PR #42908
  • Let callbacks filter the model listing routes per caller - PR #43027

CLI and coding agents​

  • Add --validate_config dry-run flag - PR #41705
  • Preserve newer installed status lines during setup - PR #42356
  • Import proxy_server once on script-style boot - PR #42584

Deployment​

  • Render a fixed replicaCount on componentized deployments when HPA is disabled - PR #42207
  • Add a quickstart compose file served from the product repo - PR #42326
  • Expose /api/event_logging/batch on the gateway allowlist - PR #42572
  • Bump wolfi-base digest to pick up glibc 2.44-r6 - PR #42643

Terraform​

  • Expose server_metadata on litellm_key so undeclared metadata is visible - PR #42453
  • Add display_name to litellm_model resource and model data sources - PR #42987

AI Integrations​

Guardrails​

  • Scan each choice's tool-call arguments apart on n>1 streams and log why a rewrite was discarded - PR #40986
  • Straiker guardrail speaks the v3 platform API (/api/v3/detect) - PR #41880
  • Add default fallback policy attachments - PR #42119
  • Opt-in include_guardrail_response returns guardrail_information in the response - PR #42327
  • Mask PII in streaming /v1/messages output - PR #42351
  • Run key-attached guardrails on /v1/videos - PR #42354
  • Store the masked output in spend logs when Presidio masks the response - PR #42441
  • Keep inherited parent guardrails when a child policy condition misses - PR #42548
  • Gate disable_global_guardrails on keys and teams to proxy admins - PR #42699
  • Stream non-Anthropic raw SSE through the post_call hook unbuffered - PR #42777
  • Keep deployment labels on cache-hit post_call guardrail rejections - PR #42780
  • Mask streamed /v1/messages output when the first upstream read is a keepalive, a data-less ping, or a split utf8 character - PR #43023

Logging and observability​

  • Migrate the sdk callback to langfuse v4. The Docker image now ships langfuse>=4.7,<5. If you install the SDK yourself, upgrade from langfuse 2.x or 3.x, which the logger now rejects at startup. Traces go through Langfuse's OpenTelemetry (OTLP) ingestion, so a self-hosted Langfuse server must accept OTLP, and trace ids and span nesting differ from the v2 integration - PR #36741
  • Replace colons in generated log filenames - PR #40452
  • Attribute provider and model_info on pre_call_hook rejections - PR #41077
  • Bound concurrent S3 uploads per flush and add opt-in JSONL batch files - PR #41258
  • Keep intercepted searches under the parent request's session and trace - PR #41711
  • Add normalized_error cluster key to error_information - PR #41715
  • Honor SSL_CERT_FILE and ssl_verify in OTLP HTTP exporters - PR #42106
  • Apply DB-stored callback redaction settings before logger init - PR #42122
  • Map OCR page markdown onto the generation output - PR #42267
  • Deliver every distinct alert queued in one flush window - PR #42314
  • Make flush() survive an event loop change - PR #42355
  • Per-team success and error sampling rates for the Arize AX callback - PR #42383
  • Price terminal Responses stream events from their inner response - PR #42385
  • Map completions, images, speech, transcription and moderation output onto the Langfuse generation output - PR #42394
  • Json.dumps with default=str so non-serializable metadata does not crash batch flush - PR #42424
  • Record the GenAI exception event through the Logs API on both OpenTelemetry lines - PR #42431
  • Map rerank and search output and the OCR, image edit and search input onto the Langfuse generation - PR #42444
  • Keep text completion choice fields beside the synthesized message - PR #42537
  • Replay approved pre-call guardrail snapshots - PR #42774
  • Scan the exceeded budget wording linearly so a crafted error message cannot stall the proxy - PR #42778
  • Pass provider response headers to callbacks on every endpoint - PR #42824
  • Root post-response service spans in their own trace linked to the request - PR #42826
  • Add model_group label to deployment request and rate limit metrics - PR #42966
  • Upload fresh events first, drop terminal failures and hour-old retries by default, opt-in adaptive concurrency - PR #43022

Secret Managers​

  • Restore secret scheduled for deletion instead of failing CreateSecret - PR #42454
  • Route secret resolution through native Rust backends - PR #42619

Spend Tracking, Budgets and Rate Limiting​

Cost tracking​

  • Honor deployment pricing for image generation - PR #39311
  • Bill batch prompts above 272K at OpenAI's long-context batch tier - PR #39861
  • Attribute CLI session spend to the per-user cli-session alias instead of the hashed session token - PR #40541
  • Warn and count $0 cost on billable requests - PR #42345
  • Honor per-second custom pricing on chat completions for every provider - PR #42403
  • Return 400 from /spend/calculate for a model with no pricing row - PR #42497
  • Capture-rate check of LiteLLM spend against the OpenAI bill - PR #43044
  • Apply a deployment's pricing override to realtime sessions - PR #43114

Budgets​

  • Renew budget reservation counter TTL while the request is in flight - PR #40322
  • Enforce virtual key budgets for JEV test routing - PR #41879
  • Return 422 instead of 429 for BudgetExceededError - PR #42097
  • Release unclaimed budget reservations at request end - PR #42304
  • Await budget redis pipeline before sync reads (internal copy of #32618) - PR #43125

Rate limiting​

  • Enforce tpm_limit and rpm_limit set on tag objects - PR #41807
  • Share model rate-limit buckets between a model_group_alias and its target - PR #42516

MCP Gateway​

MCP Gateway​

  • Handle split UTF-8 routing previews - PR #34919
  • Paginate prompt and resource discovery - PR #39189
  • Keep oauth scopes in admin api credential redaction - PR #39805
  • Keep server lists stable across refreshes - PR #41074
  • Apply post-call rewrites without stale structured output - PR #41530
  • Tools/call no longer 404s on a worker that has not served tools/list - PR #42072
  • Explain missing public client dependencies - PR #42148
  • Keep config-defined servers read-only - PR #42299
  • Restore legacy SSE and bounded cancellation cleanup - PR #42382
  • Cap an agent key's tools at what the invoking user and team may call - PR #42478
  • Preserve discovery attribution and sanitize logging headers - PR #42541
  • Preserve credential authority in DCR bridge authentication - PR #42563
  • Reject origins outside the configured allowlist - PR #42649
  • Return 401 challenge for REST token-exchange tool calls without a subject token - PR #42782
  • Forward caller bearer on REST oauth_delegate tool calls - PR #42787
  • Keep tool attribution on guardrail-blocked REST calls - PR #42790
  • Reject duplicate MCP server names and aliases - PR #42791
  • Allow ["*"] wildcard in mcp_tool_permissions to grant all current and future tools - PR #43108

Performance / Loadbalancing / Reliability improvements​

Auto Router and model routing​

  • Add configurable provider affinity header mapping - PR #41033
  • Make context-window escalation opt-in - PR #41872
  • Add JEV classifier alongside LLM classifier - PR #41886
  • Add NotFoundErrorRetries so a retry policy can pin 404 retries - PR #42045
  • Skip cooldown for background response cost poll 404s - PR #42046
  • Stamp model_group when retrieving a batch, so batch tokens are attributable (internal copy of #38499) - PR #42062
  • Native compact-to-fit across conversation APIs - PR #42074
  • Keep prompt caching affinity when the breakpoint moves - PR #42080
  • Configure heuristic v2 success threshold - PR #42252
  • Walk every entry of a fallback list after a mid-stream failure - PR #42283
  • Add group-scoped priority routing strategy - PR #42378
  • Time-windowed team reservation of deployments via model_info.access_windows - PR #42398
  • Explain fallback outcome in plain words in the raised error - PR #42509
  • Serve Responses turns from a sibling when the encrypted content origin has no boundary peer - PR #43015
  • Match provider-prefixed fallback keys for bare model groups served by wildcard deployments - PR #43062
  • Honor disable_fallbacks on mid-stream fallback - PR #43111

Caching, database and runtime​

  • Authenticate sync clusters with IAM credential providers - PR #40204
  • Coordinate v2 migration startup and qualify container recovery - PR #40932
  • Record aborted outcome when spend-log cleanup is cancelled at shutdown - PR #41213
  • Wait for the spend-log table before creating startup views - PR #41974
  • Stop re-sending un-resendable spend batches from the Redis buffer - PR #41994
  • Park requeued spend logs in Redis so they survive a pod restart during a DB outage - PR #42022
  • Default to the v2 migration resolver - PR #42105
  • Publish auth cache invalidations in the background so a wedged coordination Redis cannot stall user updates - PR #42534
  • Honor DATABASE_DISABLE_PREPARED_STATEMENTS in the litellm CLI - PR #42556
  • Keep embedding cache hits aligned with request inputs - PR #42571
  • Keep the in-flight daily spend batch when shutdown cancels the flush - PR #42593
  • Fail parked DB lookups at a deadline and flip readiness while they stall - PR #42654
  • Do not requeue a daily spend batch whose commit already left for postgres - PR #42786
  • Apply user_api_key_cache_max_size to the key object partition - PR #42796
  • Stamp provider on sync cache-hit logs so responses spend logs record provider - PR #42830

Native Rust runtime (opt-in)​

  • Preserve Python defaults with opt-in Rust dispatch - PR #42174
  • Add native disk cache backend - PR #42311
  • Add native S3 cache backend - PR #42313
  • Add native Valkey semantic cache backend - PR #42316
  • Serve RedisClusterCache natively as a Redis topology - PR #42317
  • Serve Redis Semantic caches natively in Rust - PR #42319
  • Native Azure Blob response cache backend - PR #42321
  • Serve QdrantSemanticCache natively from Rust - PR #42324
  • Add native GCS object-store cache backend - PR #42325
  • Align the cache crates with Python and wire every native backend - PR #42530
  • Dispatch Python logging through the Rust diagnostics processor - PR #42616
  • Stamp x-litellm-rust on native sync and async streams at the bridge boundary - PR #42758
  • Shape Anthropic Messages requests natively - PR #42982
  • Read secrets through Python from Rust routes and declare Rust-only routes with NO_PYTHON - PR #43057
  • Hand upstream response headers to the native Messages stream - PR #43178
  • Traverse and release the retained headers dict - PR #43274

Documentation Updates​

Documentation​

  • Stop advertising sk-1234 as the master key in shipped configs and examples - PR #42011

Tests, CI and Internal Changes​

These 184 PRs change tests, CI, contributor tooling, release packaging, or Rust runtime scaffolding that is not yet wired into a user-facing path. They do not change proxy or SDK behavior on their own

Tests (108)
  • Replace custom endpoints_client with provider SDK clients - PR #34358
  • Assert cache-priced vertex grok rows advertise supports_prompt_caching - PR #41526
  • Cover actor edges and wildcard models - PR #41769
  • Cover dashboard form journeys - PR #41773
  • Cover chat and responses registry gaps - PR #41794
  • Cover legacy lowest TPM selection - PR #41795
  • Endpoint, breakdown component and failure support in the cost harness - PR #41999
  • Embeddings, rerank, completions and moderations cost cases - PR #42020
  • Audio, image and per-unit cost cases - PR #42024
  • Passthrough route cost cases - PR #42028
  • Pricing dimension and provider reported cost cases - PR #42035
  • Restore scoped execution and credential isolation regressions - PR #42050
  • Restore MCP OAuth happy-path coverage (LIT-3467) - PR #42051
  • Provider wire cost cases - PR #42052
  • Proxy behaviour cost cases - PR #42060
  • Add autorouter estimate keys to the GCS pub/sub spend-log golden - PR #42061
  • Batch and realtime cost cases - PR #42066
  • Migrate phase 6 provider unit tests to tests/unit - PR #42107
  • Migrate wave 1 phase 2 legacy unit tests to tests/unit - PR #42108
  • Migrate phase 5 provider unit tests to tests/unit - PR #42109
  • Migrate bedrock, baseten and base_llm batch tests to tests/unit - PR #42110
  • Migrate wave 1 phase 8 legacy llm tests to tests/unit - PR #42112
  • Block external sockets at import time and add a socket policy regression test - PR #42113
  • Migrate phase 7 provider unit tests to tests/unit - PR #42114
  • Migrate nvidia, oci, ocr, oobabooga and openai legacy tests to tests/unit - PR #42115
  • Migrate phase 9 legacy llm provider tests to tests/unit - PR #42117
  • Migrate wave 1 phase 3 anthropic, apiserpent, azure and azure_ai legacy tests - PR #42118
  • Boot the unified Google proxy fixture with a real master key - PR #42120
  • Migrate wave 1 phase 1 legacy tests to tests/unit - PR #42123
  • Deflake fuzzy picker, breached-password HIBP, and MCP stdio timeout tests (rolling deflake 2026-09-22) - PR #42125
  • Migrate openai, openai_like and openrouter legacy tests to tests/unit - PR #42128
  • Migrate phase 16 legacy tests to tests/unit - PR #42131
  • Migrate phase 14 wave 2 provider tests to tests/unit - PR #42132
  • Make every tests/unit directory a package so pytest collection is unique - PR #42135
  • Migrate phase 15 legacy tests to tests/unit - PR #42136
  • Migrate phase 12 legacy llm provider tests to tests/unit - PR #42137
  • Migrate legacy provider tests to tests/unit (wave 2, phase 13) - PR #42145
  • Migrate a2a_protocol legacy tests to tests/unit (wave 3 phase 17) - PR #42160
  • Cover the release-to-release upgrade path - PR #42294
  • Fix stale budget-status and bad-database-url assertions - PR #42339
  • Add conversational matrix across chat, messages and responses - PR #42359
  • Isolate the agent read-through singleton between unknown-agent tests - PR #42389
  • Move Xiaomi MiMo coverage from live e2e to the providers wire shard - PR #42395
  • Chain a proxy-issued previous_response_id in the cost suite - PR #42396
  • Check the MCP Tools tab against the upstream's own tools/list - PR #42397
  • Make bedrock collector and secret scan timing tests deterministic - PR #42405
  • Assert a cooldown reaches a sibling replica within the 1s Redis read interval - PR #42422
  • Make the detailed-timing receive-anchor test timezone independent - PR #42429
  • Give every ui settings endpoint test a fresh settings store - PR #42430
  • Add secret manager lanes for HashiCorp Vault and CyberArk Conjur - PR #42503
  • Run the memory cell alone on the shared stack - PR #42518
  • Move unified_google_tests to gemini-3.5-flash-lite - PR #42520
  • Drop route dispatch assertions, test the bridge directly - PR #42536
  • One request lands the same spend on every surface - PR #42540
  • Guard leaked cassette patches and make injected-transport embedding tests immune - PR #42542
  • Hold every worker under an idle RSS budget before any traffic - PR #42552
  • Count a zombie grandchild as gone in the migrate deploy timeout test - PR #42570
  • Make two proxy-infra tests independent of sibling-test state - PR #42581
  • Point unit tests at model ids still in the cost map - PR #42606
  • Accept the per-size image cost keys in the price-map schema check - PR #42612
  • Point image-generation deployment price test at a live gemini row - PR #42615
  • Point CircleCI-only suites at models still in the cost map - PR #42617
  • Regression tests for August provider translation and streaming bugs - PR #42621
  • Regression tests for August cost tracking and budgeting bugs - PR #42622
  • Drop legacy InvalidStatusCode tests and pin websockets imports - PR #42624
  • Tolerate provider-side flakes on five full-suite cells - PR #42628
  • Repoint the Azure image cost test at gpt-image-2 - PR #42631
  • Raise the post-success hook error from a guardrail in the failure-hook regression - PR #42646
  • Allow skipped nodes and drop the shard cap - PR #42687
  • Add read-replica routing harness to the CircleCI integration suite - PR #42692
  • Regression tests for July provider translation, routing and streaming bugs - PR #42693
  • Regression tests for July cost tracking, budgeting and spend bugs - PR #42694
  • Add MCP gateway coverage wave 1 with a dedicated mcp shard and proxy coverage artifact - PR #42711
  • Model the blocking OCR hook as a guardrail so its raise propagates - PR #42775
  • Deterministic integration audit of the v3 platform relay - PR #42781
  • Cover customer reported cache key, cache_control, bedrock request id, responses schema, scim and tag budget contracts - PR #42785
  • Assert /v1/responses usage reports Anthropic system cache write then read - PR #42855
  • Assert /v1/models reports max_input_tokens and max_output_tokens - PR #42858
  • Cover per key tag rpm limits and tag budget_duration resets - PR #42859
  • Edge-case matrices for malformed token limits and callback_settings shapes - PR #42895
  • Take keys out of the legacy proxy, enterprise and mcp unit tests before moving them - PR #42901
  • Assert the images sent to Ollama instead of echoing them through response - PR #42905
  • Stop a comprehension variable from shadowing the body() helper - PR #42906
  • Run the files peak-memory guards without coverage tracing - PR #42914
  • Patch create_mcp_server_if_identifier_free in the store-model-in-db MCP tests - PR #42916
  • Run the cache-hit redis outage test on the shared owned_redis helper - PR #42925
  • Give the logout specs their own admin session - PR #42930
  • Accept the otel cost write as a linked root trace - PR #42931
  • Hide the LiteAdmin button in the shared admin session - PR #43033
  • Pin end-user and tag attribution from Codex-style headers on /v1/responses - PR #43093
  • Cover Azure code_interpreter container files by native id with a service-account key - PR #43122
  • Reorganize core crate tests and split cache and OCR suites - PR #43177
  • Agent clients for Claude Code, Codex and opencode - PR #43181
  • Move tests/test_litellm root and small trees into tests/unit - PR #43186
  • Move tests/test_litellm/llms into tests/unit/llms - PR #43191
  • Move tests/test_litellm integrations and secret_managers into tests/unit - PR #43194
  • Move tests/test_litellm core utils, routing, responses, caching and rust_bridge into tests/unit - PR #43199
  • Drop the repeated UNIT_FLAG key in test_unit_shard_missing_paths - PR #43212
  • Stop CI tests from downloading tokenizer files and images - PR #43257
  • Fix stale and state-leaking tests red on scheduled CircleCI - PR #43266
  • Finish the non-proxy half of tests/test_litellm - PR #43281
  • Port langfuse callbacks-in-db coverage to the local harness - PR #43282
  • Run the Langfuse DB-callback test on its own scratch database - PR #43288
  • Scope the management proxy fixture to its package so its spend monitor cannot race the spend tests - PR #43302
  • Assert only litellm-owned batch behavior and move the blank S3 env pin to an integration test - PR #43321
  • Report batch cleanup leftovers as a plain UserWarning on rc/1.104.0 - PR #43406
  • Skip the LangSmith batch serialization e2e on rc/1.104.0 until CI has a LangSmith key - PR #43418
  • Stop test modules from putting their own directory on sys.path on rc/1.104.0 - PR #43420
CI (27)
  • Wire tests/unit into CircleCI and keep draining GHA shards green - PR #42103
  • Excuse retired test-quality rules in the budget ratchet - PR #42116
  • Fix the stage-mirror batch reds and keep a redacted pytest log - PR #42143
  • Let the install smoke test boot its key-less proxy config - PR #42296
  • Skip cost map file checks on PRs that leave the cost map untouched - PR #42406
  • Print add-mask lines only under GitHub Actions - PR #42423
  • Allowlist _render_json in the recursive detector - PR #42442
  • Route credential, cost map, and UI login calls to the control plane - PR #42506
  • Drop dead misc shard paths and skip missing paths with a warning - PR #42603
  • Move the compat-matrix populator from a GCE VM to a Render cron job - PR #42608
  • Close open pull requests superseded by a merged fix on their linked issue - PR #42609
  • Remove the unused create-release workflow - PR #42696
  • Add merge smoke checks workflow with loopback-only harness and 11 curated cases - PR #42709
  • Drop litellm_internal_staging and litellm_oss_staging references, main is the only trunk - PR #42745
  • Add tests-only CircleCI pipeline with coverage and docs validation - PR #42773
  • Fix the litellm-tests unit job (sysmon, codecov on failure, env -i allowlist, selection errors, reruns param) - PR #42900
  • Move caching, proxy-extras, gateway and enterprise tests into tests/unit and run them from litellm-tests - PR #42902
  • Move tests/proxy_unit_tests to tests/unit/proxy and run the proxy-db shards from litellm-tests - PR #42903
  • Move provider-independent MCP tests into tests/unit and run mcp-integration from litellm-tests - PR #42904
  • Resolve and install the Claude Code CLI per run - PR #43038
  • Skip unpublished npm versions in the Claude Code PR-gate resolver - PR #43053
  • Run the claude_code harness unit-test trees in the lint job - PR #43077
  • Cut rc/<X.Y.0> off main every Friday at 3am Pacific - PR #43121
  • Run migrated unit selections on every event in legacy GHA shards - PR #43182
  • Drop main and litellm_* branch filters from the CircleCI litellm-main workflows - PR #43272
  • Stop stale CI reds, keep unit tests off the host env, retry CyberArk policy conflicts - PR #43294
  • Cut CircleCI wall time without loosening test isolation - PR #43347
Code quality and contributor tooling (16)
  • Remove 1,173 Any errors across 169 backend files - PR #40251
  • Define the tier contract for unit, integration and e2e - PR #42099
  • Replace Any with proven types in 30 files - PR #42127
  • Replace Any with proven types in 32 files - PR #42220
  • Extract explicit operation context and dispatch - PR #42292
  • Cap comprehensions at one for and one if clause (LIT014) - PR #42650
  • Clear fresh tech debt from the last 24 hours (rolling, 2026-09-06 to 2026-09-24) - PR #42710
  • Replace Any with proven types in 5 files - PR #42722
  • Add LIT013 flagging *-ok suppressions that suppress nothing and remove the 240 stale ones - PR #42793
  • Drop empty sections from the PR body and tighten the User Flow - PR #42794
  • Simplify pull request template into plain English questions - PR #42813
  • Restore the full pull request template (reverts #42813) - PR #42828
  • Declare litellm-owned kwargs as typed objects and derive the lists from their fields - PR #42843
  • Replace Any with proven types in 13 files - PR #42937
  • Require UI before/after screenshots and intentional UX change note in PR template - PR #43021
  • Carve harness tests out of the no-unit-tests hard rule - PR #43076
Rust runtime internals (29)
  • Split token counter backends - PR #42165
  • Add typed secret managers and shared auth adapters - PR #42173
  • Scaffold cache foundation for Python parity - PR #42196
  • Preserve Python settings semantics at the native boundary - PR #42300
  • Add CyberArk Conjur secret manager backend - PR #42303
  • Add HashiCorp Vault secret manager crate - PR #42308
  • Add Azure Key Vault secret manager backend - PR #42309
  • Add cache and secret migration foundations - PR #42328
  • Declare _CacheTestHandle.valkey_semantic in native stub - PR #42364
  • Keep native Redis semantic binding and Qdrant batch writes after merge - PR #42379
  • Align secret manager operation contexts - PR #42480
  • Add python-compat crate for Python data formats - PR #42510
  • Add standalone cost calculator - PR #42604
  • Add immutable model catalog crate - PR #42605
  • Add a guarded native response-cache resolver foundation - PR #42769
  • Add native dispatch foundation - PR #42799
  • Extend native dispatch foundation to chat completions, responses, and messages - PR #42805
  • Declare above_32k cost fields on ModelInfo - PR #42856
  • Move tests.rs files inline or under tests/ and drop autotests = false - PR #43028
  • Extract the host coroutine into its own crate - PR #43129
  • Add Rust registry validation - PR #43136
  • Replace Framer trait with tokio-util codecs - PR #43193
  • Hand out an owned Client and route all providers through the pool - PR #43245
  • Move credential inheritance and SDK limits into a driver preflight - PR #43259
  • Preserve nested optional import failures - PR #43265
  • Promote anthropic messages out of experimental_pass_through - PR #43269
  • Prepare inference and auth foundations for the gateway - PR #43287
  • Add config, router and gateway crates - PR #43289
  • Expand logging and test coverage across gateway and Anthropic messages - PR #43295
Release and packaging (4)
  • Bump litellm-enterprise 0.1.69 -> 0.1.70, litellm-proxy-extras 0.4.100 -> 0.4.101, litellm 1.103.0 -> 1.104.0 - PR #42633
  • Bump litellm-enterprise 0.1.70 -> 0.1.71, litellm-proxy-extras 0.4.101 -> 0.4.102 - PR #43120
  • Rebuild Admin UI bundle for rc/1.104.0 - PR #43372
  • Remove the top-N key cap, its follow-ups and the daily global spend rollup code from rc/1.104.0 - PR #43385

PR roll-up by ownership area​

Customer-facing PRs: 447. Tests, CI and internal PRs: 184. Total: 631

  • Models & Providers: 236
  • Other (tests, CI, internal): 184
  • Performance: 47
  • Auth & Management: 43
  • LLM API Endpoints: 25
  • Logging: 24
  • UI: 24
  • MCP: 18
  • Spend / Budgets / Rate Limits: 15
  • Guardrails: 12
  • Secret Managers: 2
  • Docs: 1

New Contributors​

Full Changelog​

Compare release contents on GitHub