1.105.0rc1 - Claude Sonnet 5.5, GPT-6.1 Sol, litellm.agent(), Lens & Agent Traces
Deploy this version​
- Docker
- Pip
docker run \
-e LITELLM_MASTER_KEY=sk-<paste-a-long-random-key> \
-e DATABASE_URL=postgresql://<user>:<password>@<host>:5432/<dbname> \
-e STORE_MODEL_IN_DB=True \
-p 4000:4000 \
docker.litellm.ai/berriai/litellm:1.105.0-rc.1
pip install litellm==1.105.0rc1
The published GitHub tag is v1.105.0-rc.1. These notes compare it with v1.104.0-rc.1, the previous release candidate cut from main. Changes backported onto rc/1.104.0 and already shipped in v1.104.0 are omitted
Customer-facing changes come first. Test, CI and internal changes are listed at the bottom
These callouts cover user-facing behavior that differs from v1.104.0, the latest stable release
ENFORCE_PRISMA_MIGRATION_CHECK=false is now ignored, so a proxy whose migrations fail will not start. This keeps a new version from serving traffic against an outdated database schema and failing requests. See the v1.104.0 breaking changes and PR #44141
Bedrock GPT-5.6, GPT-6 and GPT-6.1 move from Converse to native Chat Completions, so response ids, service_tier and reasoning fields change shape. Use bedrock/converse/<model> to stay on Converse. See PR #44307
/sso/debug/* returns 404 unless ENABLE_SSO_DEBUG=true. See PR #43150
Key Highlights​
- New models on day one: Claude Sonnet 5.5 across Anthropic, Bedrock, Bedrock Mantle, Vertex AI, Azure AI, OpenRouter and Perplexity, GPT-6.1 Sol across OpenAI, Azure, Bedrock and OpenRouter, and Grok 4.7 on Bedrock and Vertex AI, among 78 new catalog entries
- New providers and routes: Prism, Sail and Cortecs providers, native xAI batches and files, Fireworks router models, and Bedrock GPT-5.6+ served on native Chat Completions
litellm.agent(): run Claude Code, Codex, OpenCode and Deep Agents through the AI gateway from the SDK- Lens and Agent Traces: OTLP trace ingestion stored in ClickHouse with matched spend, scoped SQL over traces, a chat-style run view, and Lens (Beta) investigations run by a separate worker
- Agent identities: register agents with Entra ID identities, authenticate delegated requests and enforce authoritative agent permissions
- Faster proxy at scale: one request-scoped Redis pipeline for auth, spend, rate-limit and routing reads, gzip for buffered responses, and usage pages that page keys from the server instead of loading every key into the browser
New Providers and Endpoints​
New Providers (3 new providers)​
| Provider | Supported LiteLLM Endpoints | Description |
|---|---|---|
| Prism | /v1/chat/completions, /v1/responses, /v1/messages | Prism inference as an OpenAI-compatible provider, with DeepSeek V4 Flash and V4.1 Flash in the catalog |
| Sail | /v1/chat/completions, /v1/responses, /v1/messages | 12 Sail models, with service_tier mapped to Sail's completion window and billed at that window's price, including a new balanced tier |
| Cortecs | /v1/chat/completions, /v1/responses, /v1/messages | Cortecs, an EU LLM router, as an OpenAI-compatible provider |
Expanded provider endpoint support​
| Provider | Endpoint | What you can do |
|---|---|---|
| xAI | /v1/files, /v1/batches | Run native xAI batches and files |
| Amazon Bedrock | /v1/chat/completions | Serve GPT-5.6, GPT-6 and GPT-6.1 on bedrock-runtime's native Chat Completions, with a chat_completions/ opt-in for gpt-oss and Grok |
| Fireworks AI | /v1/chat/completions | Route to and list the auto, auto-instant and firerouter routers |
New Models / Updated Models​
New Model Support (78 new models)​
Counts represent new catalog identifiers compared with v1.104.0, including aliases and regional variants. Prices below are the values bundled in this release, in USD; runtime pricing-map reloads can update them. Input and output columns show base token rates; long-context, cache, image-token, and other specialized rates depend on the model
| Provider | Model | Context Window | Input ($/1M tokens) | Output ($/1M tokens) | Features / special pricing |
|---|---|---|---|---|---|
| Amazon Bedrock | anthropic.claude-sonnet-5-5 | 1,000,000 | $2 | $10 | Chat; Reasoning; Vision; Tool calling; Prompt caching; Structured output; PDF input; Computer use |
| Amazon Bedrock | apac.anthropic.claude-sonnet-5-5 | 1,000,000 | $2.2 | $11 | Chat; Reasoning; Vision; Tool calling; Prompt caching; Structured output; PDF input; Computer use |
| Amazon Bedrock | au.anthropic.claude-sonnet-5-5 | 1,000,000 | $2.2 | $11 | Chat; Reasoning; Vision; Tool calling; Prompt caching; Structured output; PDF input; Computer use |
| Amazon Bedrock | bedrock/us-gov-east-1/anthropic.claude-sonnet-5-5 | 1,000,000 | $2.4 | $12 | Chat; Reasoning; Vision; Tool calling; Prompt caching; Structured output; PDF input; Computer use |
| Amazon Bedrock | bedrock/us-gov-west-1/anthropic.claude-sonnet-5-5 | 1,000,000 | $2.4 | $12 | Chat; Reasoning; Vision; Tool calling; Prompt caching; Structured output; PDF input; Computer use |
| Amazon Bedrock | eu.anthropic.claude-sonnet-5-5 | 1,000,000 | $2.2 | $11 | Chat; Reasoning; Vision; Tool calling; Prompt caching; Structured output; PDF input; Computer use |
| Amazon Bedrock | global.anthropic.claude-sonnet-5-5 | 1,000,000 | $2 | $10 | Chat; Reasoning; Vision; Tool calling; Prompt caching; Structured output; PDF input; Computer use |
| Amazon Bedrock | global.openai.gpt-6.1-sol | 1,050,000 | $2 | $10 | Chat; Reasoning; Vision; Tool calling; Prompt caching |
| Amazon Bedrock | global.xai.grok-4.7 | 500,000 | $2 | $6 | Chat; Reasoning; Vision; Tool calling |
| Amazon Bedrock | jp.anthropic.claude-sonnet-5-5 | 1,000,000 | $2.2 | $11 | Chat; Reasoning; Vision; Tool calling; Prompt caching; Structured output; PDF input; Computer use |
| Amazon Bedrock | openai.gpt-6.1-sol | 1,050,000 | $2 | $10 | Chat; Reasoning; Vision; Tool calling; Prompt caching |
| Amazon Bedrock | us-gov.anthropic.claude-sonnet-5-5 | 1,000,000 | $2.4 | $12 | Chat; Reasoning; Vision; Tool calling; Prompt caching; Structured output; PDF input; Computer use |
| Amazon Bedrock | us.anthropic.claude-sonnet-5-5 | 1,000,000 | $2.2 | $11 | Chat; Reasoning; Vision; Tool calling; Prompt caching; Structured output; PDF input; Computer use |
| Amazon Bedrock | us.openai.gpt-6.1-sol | 1,050,000 | $2.2 | $11 | Chat; Reasoning; Vision; Tool calling; Prompt caching |
| Amazon Bedrock | us.xai.grok-4.7 | 500,000 | $2.2 | $6.6 | Chat; Reasoning; Vision; Tool calling |
| Amazon Bedrock | xai.grok-4.7 | 500,000 | $2 | $6 | Chat; Reasoning; Vision; Tool calling |
| Anthropic | claude-sonnet-5-5 | 1,000,000 | $2 | $10 | Chat; Reasoning; Vision; Tool calling; Prompt caching; Structured output; PDF input; Web search; Computer use |
| Azure AI | azure_ai/MAI-Cyber-1-Flash | 256,000 | $0.6 | $3.5 | Chat; Reasoning; Tool calling; Prompt caching |
| Azure AI | azure_ai/claude-sonnet-5-5 | 1,000,000 | $2 | $10 | Chat; Reasoning; Vision; Tool calling; Prompt caching; Structured output; PDF input; Computer use |
| Azure AI | azure_ai/gpt-6.1-sol | 922,000 | $2 | $10 | Chat; Reasoning; Vision; Tool calling; Prompt caching; Structured output; PDF input; Web search; Computer use |
| Azure AI | azure_ai/mistral-ocr-2505 | - | - | - | OCR; ocr_cost_per_page: $0.001 |
| Azure AI | azure_ai/mistral-ocr-2512 | - | - | - | OCR; ocr_cost_per_page: $0.002 |
| Azure OpenAI | azure/gpt-6.1-sol | 922,000 | $2 | $10 | Chat; Reasoning; Vision; Tool calling; Prompt caching; Structured output; PDF input; Web search; Computer use |
| Azure OpenAI | azure/gpt-6.1-sol-2026-09-29 | 922,000 | $2 | $10 | Chat; Reasoning; Vision; Tool calling; Prompt caching; Structured output; PDF input; Web search; Computer use |
| Baseten | baseten/deepseek-ai/DeepSeek-V4.1-Flash-Fast | 1,048,576 | $0.6 | $2.4 | Chat; Reasoning; Vision; Tool calling; Prompt caching; Structured output |
| Bedrock Mantle | bedrock_mantle/anthropic.claude-opus-5-5 | 1,000,000 | $4.4 | $22 | Chat; Reasoning; Vision; Tool calling; Prompt caching; Structured output; PDF input; Computer use |
| Bedrock Mantle | bedrock_mantle/anthropic.claude-sonnet-5-5 | 1,000,000 | $2.2 | $11 | Chat; Reasoning; Vision; Tool calling; Prompt caching; Structured output; PDF input; Computer use |
| Bedrock Mantle | bedrock_mantle/openai.gpt-6.1-sol | 1,050,000 | $2.2 | $11 | Responses; Reasoning; Vision; Tool calling; Prompt caching; Structured output |
| Bedrock Mantle | bedrock_mantle/us-gov-west-1/anthropic.claude-opus-5-5 | 1,000,000 | $4.8 | $24 | Chat; Reasoning; Vision; Tool calling; Prompt caching; Structured output; PDF input; Computer use |
| Bedrock Mantle | bedrock_mantle/us-gov-west-1/anthropic.claude-sonnet-5-5 | 1,000,000 | $2.4 | $12 | Chat; Reasoning; Vision; Tool calling; Prompt caching; Structured output; PDF input; Computer use |
| Databricks | databricks/databricks-claude-opus-5-5 | 1,000,000 | $4.00001 | $19.99998 | Chat; Reasoning; Vision; Tool calling; Prompt caching |
| Fireworks AI | fireworks_ai/accounts/fireworks/routers/auto | - | - | - | Chat; Reasoning; Tool calling; Structured output |
| Fireworks AI | fireworks_ai/accounts/fireworks/routers/auto-instant | - | - | - | Chat; Reasoning; Tool calling; Structured output |
| Fireworks AI | fireworks_ai/accounts/fireworks/routers/firerouter | - | - | - | Chat; Reasoning; Tool calling; Structured output |
| Nebius | nebius/Qwen/Qwen3.8-27B | 262,144 | $0.45 | $3 | Chat; Reasoning; Tool calling |
| OpenAI | gpt-4o-mini-tts-2025-03-20 | - | $0.6 | $10 | Speech; output_cost_per_second: $0.00025 |
| OpenAI | gpt-6.1-sol | 922,000 | $2 | $10 | Chat; Reasoning; Vision; Tool calling; Prompt caching; Structured output; PDF input; Web search; Computer use |
| OpenRouter | openrouter/anthropic/claude-sonnet-5.5 | 1,000,000 | $2 | $10 | Chat; Reasoning; Vision; Tool calling; Prompt caching; Structured output; PDF input; Web search |
| OpenRouter | openrouter/anthropic/claude-sonnet-5.5:batch | 1,000,000 | $1 | $5 | Chat; Reasoning; Vision; Tool calling; Prompt caching; Structured output; PDF input; Web search |
| OpenRouter | openrouter/apodex/apodex-1.1-mini:free | 262,144 | $0 | $0 | Chat; Reasoning; Tool calling; Structured output |
| OpenRouter | openrouter/openai/gpt-6.1-sol | 1,050,000 | $2 | $10 | Chat; Reasoning; Vision; Tool calling; Prompt caching; Structured output; PDF input; Web search |
| OpenRouter | openrouter/openai/gpt-6.1-sol-pro | 1,050,000 | $2 | $10 | Chat; Reasoning; Vision; Tool calling; Prompt caching; Structured output; PDF input; Web search |
| OpenRouter | openrouter/openai/gpt-6.1-sol-pro:batch | 1,050,000 | $1 | $5 | Chat; Reasoning; Vision; Tool calling; Prompt caching; Structured output; PDF input; Web search |
| OpenRouter | openrouter/openai/gpt-6.1-sol:batch | 1,050,000 | $1 | $5 | Chat; Reasoning; Vision; Tool calling; Prompt caching; Structured output; PDF input; Web search |
| OpenRouter | openrouter/typesafe/jev-router | 1,000,000 | $0 | $0 | Chat; Reasoning; Vision; Tool calling; Structured output; PDF input; Audio input |
| OpenRouter | openrouter/unbiased/pareto-26.10-preview | 1,048,576 | $0.8 | $3.2 | Chat; Vision; Tool calling; Prompt caching |
| Perplexity | perplexity/anthropic/claude-fable-5-1 | - | $10 | $50 | Responses |
| Perplexity | perplexity/anthropic/claude-opus-5-5 | - | $4 | $20 | Responses |
| Perplexity | perplexity/anthropic/claude-sonnet-5-5 | - | $2 | $10 | Responses; Tool calling; Web search |
| Perplexity | perplexity/google/gemini-3.8-flash | - | $0.75 | $3.75 | Responses |
| Perplexity | perplexity/openai/gpt-6-luna | - | $0.1 | $0.5 | Responses |
| Perplexity | perplexity/openai/gpt-6-sol | - | $2 | $10 | Responses |
| Perplexity | perplexity/openai/gpt-6.1-sol | - | $2 | $10 | Responses |
| Perplexity | perplexity/xai/grok-4.7 | - | $2 | $6 | Responses |
| Prism | prism/deepseek-v4-flash | 1,000,000 | $0.17 | $0.21 | Chat; Reasoning; Tool calling; Prompt caching; Structured output |
| Prism | prism/deepseek-v4.1-flash | 1,000,000 | $0.17 | $0.63 | Chat; Reasoning; Vision; Tool calling; Prompt caching; Structured output |
| Sail | sail/Qwen/Qwen3.6-35B-A3B | 262,144 | $0.05 | $0.4 | Chat; Reasoning; Vision; Tool calling; Prompt caching; Structured output |
| Sail | sail/deepseek-ai/DeepSeek-V4-Flash-0731 | 1,048,576 | $0.09 | $0.18 | Chat; Reasoning; Tool calling; Prompt caching; Structured output |
| Sail | sail/deepseek-ai/DeepSeek-V4-Pro-0813 | 1,048,576 | $0.92 | $2.77 | Chat; Reasoning; Tool calling; Prompt caching; Structured output |
| Sail | sail/deepseek-ai/DeepSeek-V4.1-Flash | 1,048,576 | $0.15 | $0.6 | Chat; Reasoning; Tool calling; Prompt caching; Structured output |
| Sail | sail/google/gemma-4-12B-it | 16,384 | $0.3 | $2 | Chat; Reasoning; Tool calling; Prompt caching; Structured output |
| Sail | sail/google/gemma-4-31B-it | 256,000 | $0.4 | $0.6 | Chat; Reasoning; Vision; Tool calling; Prompt caching; Structured output |
| Sail | sail/moonshotai/Kimi-K2.6 | 262,144 | $1 | $4 | Chat; Reasoning; Vision; Tool calling; Prompt caching; Structured output |
| Sail | sail/moonshotai/Kimi-K3 | 1,048,576 | $2.5 | $12.5 | Chat; Reasoning; Tool calling; Prompt caching; Structured output |
| Sail | sail/nvidia/Gemma-4-31B-IT-NVFP4 | 262,144 | $0.14 | $0.4 | Chat; Reasoning; Vision; Tool calling; Prompt caching; Structured output |
| Sail | sail/openai/gpt-oss-120b | 131,072 | $0.06 | $0.4 | Chat; Reasoning; Tool calling; Prompt caching; Structured output |
| Sail | sail/zai-org/GLM-5.3 | 1,048,576 | $0.98 | $3.08 | Chat; Reasoning; Tool calling; Prompt caching; Structured output |
| Sail | sail/zai-org/GLM-5.3-Flash | 1,048,576 | $0.11 | $0.35 | Chat; Reasoning; Tool calling; Prompt caching; Structured output |
| Together AI | together_ai/Salesforce/Llama-Rank-V1 | 8,192 | $0.1 | $0 | Rerank |
| Together AI | together_ai/meta-llama/Meta-Llama-3.1-8B | 16,384 | $0.2 | $0.2 | Completion |
| Vertex AI | vertex_ai/claude-sonnet-5-5 | 1,000,000 | $2 | $10 | Chat; Reasoning; Vision; Tool calling; Prompt caching; Structured output; PDF input; Computer use |
| Vertex AI | vertex_ai/claude-sonnet-5-5@default | 1,000,000 | $2 | $10 | Chat; Reasoning; Vision; Tool calling; Prompt caching; Structured output; PDF input; Computer use |
| Vertex AI | vertex_ai/gemini-3.8-flash-lite-tts | 8,192 | $0.5 | $6 | Speech |
| Vertex AI | vertex_ai/gemini-3.8-flash-tts | 8,192 | $0.5 | $9 | Speech |
| Vertex AI | vertex_ai/xai/grok-4.7 | 524,288 | $2 | $6 | Chat; Reasoning; Vision; Tool calling; Prompt caching; Structured output |
| Voyage AI | voyage/rerank-1 | 8,000 | $0.05 | $0 | Rerank |
| Voyage AI | voyage/rerank-lite-1 | 4,000 | $0.02 | $0 | Rerank |
| Voyage AI | voyage/voyage-large-2-instruct | 16,000 | $0.12 | $0 | Embedding |
Updated pricing (59 models)​
| Provider / model | Changed token prices (USD per 1M tokens) |
|---|---|
azure/eu/gpt-6-astra | Input: $11 to $12; Output: $55 to $60; Cache read: $1.1 to $1.2; Cache write: $13.75 to $15 |
azure/gpt-4o-mini | Input: $0.165 to $0.15; Output: $0.66 to $0.6 |
azure/gpt-4o-mini-tts | Input: $2.5 to $0.6 |
azure_ai/deepseek-v4.1-flash | Input: $0.375 to $0.3; Output: $1.5 to $1.2; Cache read: $0.008 to $0.006 |
azure_ai/grok-4.6 | Input: $2 to $1.25 |
bedrock/ap-southeast-2/minimax.minimax-m2.5 | Input: $0.309 to $0.31; Output: $1.236 to $1.24 |
fireworks-ai-up-to-4b | Input: $0.2 to $0.1; Output: $0.2 to $0.1 |
fireworks_ai/accounts/fireworks/models/deepseek-v4p1-flash | Input: $0.22 to $0.3; Output: $0.66 to $1.2; Cache read: $0.007 to $0.006 |
fireworks_ai/deepseek-v4p1-flash | Input: $0.22 to $0.3; Output: $0.66 to $1.2; Cache read: $0.007 to $0.006 |
github_copilot/claude-haiku-4.5 | Input: not set to $1; Output: not set to $5; Cache read: not set to $0.1; Cache write: not set to $1.25 |
github_copilot/claude-sonnet-4 | Input: not set to $3; Output: not set to $15; Cache read: not set to $0.3; Cache write: not set to $3.75 |
github_copilot/gpt-5-mini | Input: not set to $0.25; Output: not set to $2; Cache read: not set to $0.025 |
github_copilot/gpt-5.3-codex | Input: not set to $1.75; Output: not set to $14; Cache read: not set to $0.175 |
meta.llama3-1-405b-instruct-v1:0 | Input: $5.32 to $2.4; Output: $16 to $2.4 |
meta.llama3-1-70b-instruct-v1:0 | Input: $0.99 to $0.72; Output: $0.99 to $0.72 |
meta.llama3-2-11b-instruct-v1:0 | Input: $0.35 to $0.16; Output: $0.35 to $0.16 |
meta.llama3-2-90b-instruct-v1:0 | Input: $2 to $0.72; Output: $2 to $0.72 |
mistral.mistral-large-2407-v1:0 | Input: $3 to $2; Output: $9 to $6 |
moonshotai.kimi-k3 | Input: $3 to $3.3; Output: $15 to $16.5; Cache read: $0.3 to $0.33; Cache write: $3.75 to $4.125 |
openrouter/deepseek/deepseek-chat | Input: $0.32 to $0.2574; Output: $0.89 to $1.0287 |
openrouter/deepseek/deepseek-chat-v3-0324 | Input: $0.25 to $0.29; Output: $1 to $1.14; Cache read: not set to $0.11 |
openrouter/deepseek/deepseek-v3.1-terminus | Input: $0.27 to $0.3 |
openrouter/deepseek/deepseek-v3.2 | Input: $0.269 to $0.28; Output: $0.4 to $0.42; Cache read: $0.1345 to $0.028 |
openrouter/deepseek/deepseek-v4-flash | Input: $0.049 to $0.04186; Output: $0.098 to $0.08372; Cache read: $0.0098 to $0.008372 |
openrouter/deepseek/deepseek-v4-flash-0731 | Input: $0.03 to $0.0108; Output: $0.32 to $1.28; Cache read: $0.016 to $0.0108 |
openrouter/deepseek/deepseek-v4-flash-vision-exp | Input: $0.22 to $0.2156; Output: $0.66 to $0.6468; Cache read: $0.007 to $0.00686 |
openrouter/deepseek/deepseek-v4-pro | Input: $0.844944 to $0.2088; Output: $1.689888 to $0.4176; Cache read: $0.070412 to $0.0174 |
openrouter/deepseek/deepseek-v4-pro-0813 | Input: $0.462 to $1.32; Output: $1.386 to $3.96; Cache read: $0.0154 to $0.044 |
openrouter/deepseek/deepseek-v4.1-flash | Input: $0.3 to $0.03; Output: $1.2 to $0.5; Cache read: $0.006 to $0.01 |
openrouter/google/gemma-4-26b-a4b-it | Input: $0.09 to $0.0765; Output: $0.3 to $0.255; Cache read: $0.05 to $0.0425 |
openrouter/inclusionai/ling-3.0-flash-vl | Input: $0.06 to $0.021; Output: $0.18 to $0.0616; Cache read: $0.012 to $0.0042 |
openrouter/meta/muse-glimmer-30b | Input: $0.3 to $0.35; Output: $1.2 to $1.5 |
openrouter/minimax/minimax-m1 | Input: $0.4 to $0.55 |
openrouter/minimax/minimax-m2.7 | Input: $0.3 to $0.21; Output: $1.2 to $0.84; Cache read: $0.06 to $0.042 |
openrouter/moonshotai/kimi-k2.6 | Input: $0.95 to $0.43415; Output: $4 to $1.828; Cache read: $0.16 to $0.07312 |
openrouter/moonshotai/kimi-k2.7-code | Input: $0.6562 to $0.6712; Output: $3.3 to $3.35 |
openrouter/moonshotai/kimi-k3 | Input: $3 to $0.4357; Output: $15 to $10; Cache read: $0.3 to $0.4357 |
openrouter/nvidia/nemotron-3.5-lightning | Input: $0.08 to $0.0595; Output: $0.2 to $0.17; Cache read: $0.04 to $0.02975 |
openrouter/openai/gpt-5.6-sol-pro | Input: $2 to $4; Output: $10 to $20; Cache read: $0.2 to $0.4; Cache write: $2.5 to $5 |
openrouter/openai/gpt-oss-120b | Input: $0.15 to $0.037; Output: $0.6 to $0.17; Cache read: $0.075 to not set |
openrouter/openai/gpt-oss-20b | Cache read: $0.03 to $0.009 |
openrouter/prism-ml/ternary-bonsai-2-27b | Cache read: not set to $0.0375 |
openrouter/qwen/qwen3-30b-a3b-instruct-2507 | Input: $0.1 to $0.04815; Output: $0.3 to $0.19305 |
openrouter/qwen/qwen3-vl-30b-a3b-instruct | Input: $0.13 to $0.15; Output: $0.52 to $0.6 |
openrouter/qwen/qwen3.5-35b-a3b | Input: $0.3125 to $0.1625; Output: $1.25 to $1.3 |
openrouter/qwen/qwen3.6-27b | Output: $2.7 to $3.2 |
openrouter/x-ai/grok-4.7 | Input: $1.6 to $2; Output: $4.8 to $6; Cache read: $0.4 to $0.5 |
openrouter/z-ai/glm-4.6v | Cache read: $0.055 to $0.05 |
openrouter/z-ai/glm-4.7 | Input: $0.4 to $0.6; Output: $1.75 to $2.2; Cache read: $0.08 to $0.11 |
openrouter/z-ai/glm-5.1 | Input: $0.966 to $0.9646; Output: $3.036 to $3.0316; Cache read: $0.1794 to $0.17914 |
openrouter/z-ai/glm-5.2 | Input: $0.6496 to $0.41; Output: $2.0416 to $3.99; Cache read: $0.12064 to $0.26 |
openrouter/z-ai/glm-5.3 | Input: $1.4 to $0.2219; Output: $4.4 to $3.39; Cache read: $0.26 to $0.1775 |
openrouter/z-ai/glm-5.3-flash | Input: $0.045 to $0.15; Output: $0.6 to $0.5; Cache read: $0.0285 to $0.03 |
openrouter/~x-ai/grok-latest | Input: $1.6 to $2; Output: $4.8 to $6; Cache read: $0.4 to $0.5 |
perplexity/openai/gpt-5.6-sol | Input: $5 to $4; Output: $30 to $20; Cache read: $0.5 to $0.4 |
us.meta.llama3-1-405b-instruct-v1:0 | Input: $5.32 to $2.4; Output: $16 to $2.4 |
us.meta.llama3-1-70b-instruct-v1:0 | Input: $0.99 to $0.72; Output: $0.99 to $0.72 |
us.meta.llama3-2-11b-instruct-v1:0 | Input: $0.35 to $0.16; Output: $0.35 to $0.16 |
us.meta.llama3-2-90b-instruct-v1:0 | Input: $2 to $0.72; Output: $2 to $0.72 |
The registry also updates capability flags, context/output limits, non-token rates, and deprecation dates on 244 more entries
Removed catalog entries (1)
azure_ai/muse-spark-1.3
New providers​
- Add Prism provider - PR #41961
- Add Sail as a provider, with service_tier mapped to its completion window - PR #42840
- Add Cortecs as an OpenAI-compatible provider - PR #43872
Amazon Bedrock​
- Surface a converse-stream 200 that decodes to no events as a 502 instead of an empty turn - PR #43213
- Keep the provider status code on unprocessable image errors - PR #43416
- Add xai grok-4.7 pricing and sync llama, mistral large 2407 and minimax m2.5 prices - PR #43623
- Add bedrock_mantle rows for claude opus 5.5 and sonnet 5.5 - PR #43647
- Add openai gpt-6.1-sol global and base rows - PR #43758
- Add openai.gpt-6.1-sol us geo cris and Mantle rows - PR #43763
- Add beta header for output config in message - PR #43778
- Set gpt-6.1-sol max output tokens to 131072 - PR #43782
- Keep applicable beta headers - PR #43829
- Add beta header for thinking display updates - PR #43832
- Add beta for mid-conversation tool changes - PR #43833
- Accept Converse messages with no content key - PR #43936
- Serve gpt-5.6+ chat completions natively by default, with chat_completions/ opt-in for gpt-oss and grok - PR #44307
Anthropic​
- Drop thinking blocks with empty thinking text, not just missing signature - PR #38049
- Stand default cache points down when extra_body hides a direct client mark - PR #43341
- Forward the per-turn-control beta to Azure AI Foundry - PR #43415
- Add claude-sonnet-5-5 model pricing - PR #43586
- Correct Claude Sonnet 5.5 capabilities and provider keys - PR #43587
- Forward the dangerous-tool-use beta to Azure AI Foundry - PR #43934
- Keep thinking display updates beta - PR #43969
Fireworks AI​
- Route and list the auto, auto-instant and firerouter routers - PR #43641
Gemini and Vertex AI​
- Stop importing the vertexai SDK in partner-model completion - PR #42274
- Keep legacy bucket_name in credential resolution and add GCS_BATCH_BUCKET_NAME env var - PR #42803
- Make Gemma fake streams work with traced Responses - PR #43147
- Forward seed to the Gemini API instead of rejecting it - PR #43197
- Consider tools when validating context caching min tokens - PR #43319
- Preserve proxy_server_request in completion adapter - PR #43536
- Forward the per-turn-control beta for per-message output_config - PR #43558
OpenAI​
- Exclude fine-tuned and custom gpt-5-chat aliases from gpt-5 reasoning path - PR #43185
- Add openai gpt-6.1-sol from the pricing page - PR #43738
- Add openai gpt-6-astra ultrafast tier prices from the pricing page - PR #43745
xAI​
- Add native xAI batches and files support - PR #42812
Hosted vLLM​
- Keep reasoning_content on replayed assistant messages - PR #43599
Model catalog and pricing​
- Sync openrouter prices and add perceptron-mk1.5 - PR #43246
- Add fireworks us-only deepseek v4.1 flash priority prices - PR #43247
- Add typesafe/jev-router to the cost map - PR #43248
- Add fireworks priority prices for muse glimmer 30b and deepseek v4 flash vision exp - PR #43252
- Correct fireworks_ai deepseek-v4p1-flash pricing - PR #43253
- Registry audit 2026-09-26, MAI-Image-2.5-Flash price, Databricks Claude Opus 5.5, Azure Foundry retirement dates - PR #43254
- Remove duplicate openrouter/perceptron/perceptron-mk1.5 entry - PR #43273
- Price fireworks deepseek v4.1 flash at the prices api value - PR #43311
- Sync OpenRouter, Together, Cohere and Azure AI registry values with official sources - PR #43337
- Correct azure gpt-4o-mini tts, transcribe, alias and MAI-Image-2.5 prices - PR #43357
- Sync openrouter prices for deepseek, minimax, qwen and glm rows - PR #43384
- Drop stale cache hit field from openrouter deepseek-v4-pro-0813 - PR #43389
- Update azure_ai/grok-4.6 input price from Azure pricing page - PR #43440
- Add azure_ai/MAI-Cyber-1-Flash - PR #43446
- Sync openrouter prices from the models API - PR #43506
- Add deprecation_date to together_ai Salesforce/Llama-Rank-V1 - PR #43507
- Set together_ai gpt-oss-20b and gemma-4-31B-it deprecation_date to 2026-09-15 - PR #43509
- Add mistral ocr pricing and azure max output limits - PR #43530
- Registry audit 2026-09-28, openai deep-research shutdown dates, azure deepseek v4.1 flash direct price, vertex gemini 3.8 live avatar price, drop azure_ai/muse-spark-1.3 - PR #43566
- Add web search flag and model page source to anthropic claude-sonnet-5-5 - PR #43584
- Add tool calling and reasoning flags, correct max output for nebius DeepSeek-V4.1-Flash - PR #43588
- Take azure limits for deepseek-v4-flash-0731 and v3.2-speciale - PR #43597
- Align Azure, Bedrock, Copilot, Gemini, Groq, OpenAI and OpenRouter entries with official docs - PR #43598
- Add Vertex batch cache prices to vertex_ai/claude-sonnet-5-5 - PR #43602
- Add baseten DeepSeek-V4.1-Flash-Fast - PR #43735
- Lower fireworks up-to-4b size tier to the pricing page price - PR #43740
- Add azure and openrouter gpt-6.1-sol rows - PR #43744
- Take azure context limits from models-sold-directly - PR #43759
- Add deprecation_date to two together_ai nvidia rows - PR #43809
- Add fireworks priority prices for ember-1, nemotron and glm 5.3 us rows - PR #43811
- Add Gemini Veo, Mistral and Azure Claude 4.5 deprecation dates - PR #43857
- Add openai gpt-image-2.5 batch prices from the pricing page - PR #43869
- Add vertex_ai gemini-3.8 flash tts rows - PR #43876
- Add deprecation date for anthropic claude-sonnet-4-5 - PR #43898
- Add perplexity, openrouter, voyage and nebius models and fix registry metadata - PR #43907
- Raise baseten DeepSeek-V4.1-Flash max output to 262144 - PR #43916
- Add fireworks inkling priority prices from the prices api - PR #43949
- Sync openrouter prices from the models API - PR #43950
- Set supports_vision true on GLM-5.3-Flash - PR #43951
- Reprice fireworks deepseek v4.1 flash to the 2026-10-01 pricing update - PR #44024
- Add vertex_ai/xai/grok-4.7 pricing - PR #44059
- Restore later azure Models API retirement dates and date gpt-6.1-sol - PR #44072
- Sync openrouter prices from the models API - PR #44105
- Add azure_ai deprecation dates from the Azure retired models page - PR #44142
- Take azure_ai claude-sonnet-4-5 retirement date from the Azure schedule - PR #44145
LLM API Endpoints​
Responses API​
- Stream guardrail pre-call block as SSE with a typed output item - PR #42507
- Fall back on pre-output stream drops, fail truncated streams, honor request_timeout - PR #43133
- Record streamed /v1/responses container ownership before the response.completed frame - PR #43140
- Hold Responses lifecycle events until output so a pre-output fallback announces one response - PR #43238
- Run stream failure and success hooks on the iterating loop instead of blocking it - PR #43270
- Emit the reasoning item on streaming /v1/responses for signature-only thinking - PR #43414
Anthropic Messages API​
- Send a real error event when a /v1/messages stream fails - PR #41826
- Surface Responses bridge stream failures as Anthropic error events - PR #43126
- Rename the
litellm.llms.anthropic.experimental_pass_throughpackage topass_through- PR #43329 - Stream /v1/messages lifecycle frames live when no fallback can take over - PR #43600
Agents and Agent-to-Agent​
- Add identity storage and validation contracts - PR #43720
- Enforce authoritative agent permissions - PR #43721
- Authenticate Entra identities and delegated requests - PR #43722
- Add identity registration and dashboard controls - PR #43723
litellm.agent() (SDK)​
- Add
litellm.agent()to run Claude Code, Codex, OpenCode and Deep Agents through the AI gateway - PR #43885
Vector Stores and RAG​
- Enforce key/team vector_stores allowlist on /v1/rag/query - PR #43953
Audio​
- Honor base_url alias for Groq Whisper and report it as the api base - PR #43917
Pass-through endpoints​
- Strip caller credentials from websocket passthrough - PR #43855
- Relay Azure passthrough body model groups through the router - PR #43896
- Preserve decision request bodies under token limits - PR #43920
General​
- Keep
_litellm_*kwargs out of provider request bodies by construction - PR #43221 - Validate stream_chunk_size once, before any provider call - PR #43222
- Keep the submitted body out of 422 validation errors - PR #43231
- Salvage concatenated JSON tool call arguments - PR #43260
- Count Gemini function_declarations tools - PR #43417
- Categorize internal param appropriately to prevent leaking into request - PR #43783
- Keep silent_model out of embedding provider requests - PR #44064
Management Endpoints / UI​
Admin UI​
- Right-align money and count columns across tables - PR #37889
- Render access group MCP and agent selections as wrapping chips - PR #41228
- Surface x-litellm-call-id in Logs search, table and drawer - PR #42436
- Filter tags by name and description on the Tag Management page - PR #42949
- Rename All Models tab to Deployed Models and model filters to All Proxy Models - PR #43638
- Add model leaderboard page - PR #43649
- Show the user who owns the key that discovered a tool - PR #43892
- Adopt the new LiteLLM logo and monogram - PR #43913
- Drop the Beta badge from the Cost Optimization nav item - PR #43967
- Give model leaderboard a distinct trophy icon - PR #44036
- Show daily token totals on the model leaderboard - PR #44044
- Leave unset callback select params out of the save payload - PR #44213
- Shrink the sidebar logo so it stops outweighing page titles - PR #44247
Usage and analytics​
- Aggregated daily activity endpoints return the top 100 keys in
breakdown.api_keysby default (totals unchanged); setapi_key_limitup to 1000 - PR #43398 - Bounded daily activity routes (aggregated, search, model_top_keys, export, cache_leakage_keys) for all usage entities - PR #43408
- Usage pages consume bounded daily activity routes instead of storing all keys client-side - PR #43409
- Recover session key owners from daily spend for usage attribution - PR #43642
- Look up hashed key names with two spend log rows per key - PR #43656
- Add native ROI calculator for gateway spend vs merged PRs - PR #43669
- Attribute completed batch cost rows to /batches in daily activity - PR #43870
- Keep NULL entity ids when excluding entity ids - PR #44139
- Reject non-canonical daily activity dates - PR #44143
Keys, teams and authentication​
- Delete large teams without per-member transaction fan-out - PR #42998
- Log key owner identity on expired key auth failures - PR #43105
- Gate /sso/debug routes behind ENABLE_SSO_DEBUG, off by default - PR #43150
- Persist SSO display name as user_alias on login - PR #44065
Proxy configuration​
- Add LITELLM_DISABLE_LAZY_ROUTES to register optional routers at startup - PR #43911
CLI and coding agents​
- Reuse saved agent setup and add reconfigure - PR #43392
Deployment​
- Keep wheel paths under Windows MAX_PATH for Store Python - PR #43903
- Drop the no-op PROXY_EXTRAS_SOURCE switch from the non-root image - PR #44097
- Always exit when database setup fails at boot - PR #44141
Terraform​
- Accept 2xx status codes in unified_access_group create - PR #42461
AI Integrations​
Lens and Agent Traces​
- Add Rust storage foundation - PR #43819
- Analyze agent activity with a separate worker - PR #43889
- Agent traces tab on logs with timeline and otel setup guide - PR #43891
- Correct ClickHouse rollup partitioning, dedupe keys, and retention changes - PR #43901
- Port OTLP ingestion to current trace foundation - PR #43915
- Store spend in ClickHouse automatically - PR #43928
- Investigate sampled traces and retain batch results - PR #43942
- Agent traces open in a side drawer with a chat-style run view - PR #43972
- Improve trace ingestion and trace details - PR #43975
- Track worker spend through virtual keys - PR #43989
- Rename Lens internals and move its API from
/engineto/lens- PR #44034 - Inject tracing receiver and access context - PR #44035
- Move traces and setup into Lens - PR #44068
- Normalize agent spans in Rust - PR #44071
- Add scoped SQL queries and schema-aware help - PR #44085
- Simplify setup and investigation workflow - PR #44089
- Add test trace, tracing key and otel endpoints to tracing setup - PR #44090
- Label lens trace services as agents - PR #44116
- Recalculate ClickHouse TTL info only on retention changes - PR #44117
Guardrails​
- Honor experimental_use_latest_role_message_only on every request shape - PR #42447
- Enable explicit PANW MCP output scanning - PR #43109
- Honor litellm_params.timeout in every HTTP guardrail - PR #43134
- Fix agent 365 to the production endpoint and log the opt-in fail_open at error level - PR #43189
- Scan retrieved vector store chunks with the request's pre-call guardrails - PR #43271
- Block private destinations in custom code http_request and bound guardrail execution time - PR #43280
- Preserve Presidio output selection and restoration - PR #43401
- Scan and mask top-level instructions with guardrails - PR #43629
- Send a configured gateway_name from noma_v2 to Noma - PR #43678
- Send request conversation and tool calls to post-call monitor - PR #43770
- Scan Responses API input in Azure Prompt Shield - PR #43786
- Treat an unknown straiker api_version as unset instead of skipping the guardrail - PR #43956
- Scan Responses API input in Azure Text Moderation - PR #43965
- Route Straiker
sk_agt_keys to v3 and fail closed on a missing verdict - PR #44011 - Restore Azure guardrail get_user_prompt dispatch and allow logging - PR #44067
Logging and observability​
- Add Databricks Zerobus trace logging callback - PR #42013
- Keep the DataLakeServiceClient alive until its TTL elapses - PR #43082
- Scrub PII and secrets inside object reprs and nested locals, add SENTRY_SEND_DEFAULT_PII opt-in - PR #43123
- Keep team callback credentials out of the stored request body - PR #43217
- Redact raw_request when turn_off_message_logging is set in the proxy config - PR #43219
- Detach post-response service spans by request phase, name redis spans by operation - PR #43237
- Excluded_services opt-out for datastore spans on tenant destinations - PR #43278
- Add SigNoz preset for OpenTelemetry v2 - PR #43296
- Deliver spans to app.langtrace.ai/api/trace with x-api-key - PR #43322
- Send cache and reasoning tokens in langfuse usage_details - PR #43553
- Keep request-body credentials out of stored spend-log requests - PR #43635
- Add s3_partition_granularity option for hourly S3 folders - PR #43748
- Register a UI-configured arize callback next to otel under OTel v2 - PR #43906
- Name Data Lake objects without base64 padding or slashes - PR #43914
- Keep tool payloads and logprobs unmasked in stored spend logs - PR #44075
- Tolerate non-dict callback_settings.otel and ignore bare EXCLUDED_SERVICES env - PR #44086
- Keep client call ids from sharing one Data Lake file - PR #44099
Spend Tracking, Budgets and Rate Limiting​
Cost tracking​
- Keep the served service_tier on streamed chunks and spend rows - PR #42870
- Record in spend logs whether a request used a client-forwarded Anthropic OAuth token - PR #43063
- Bill chat per-second pricing once with a new cost_per_second field - PR #43614
- Stop copying optional_params into response hidden params - PR #43637
- Bill ultrafast prompts above 272k at the ultrafast long-context rates - PR #43764
- Bill service tiers at catalog rates for custom-priced deployments - PR #43890
Budgets​
- Count provider budget spend on every API surface - PR #38172
- Email alerts at configured percentages of a team member budget - PR #42665
- Resolve model_group_alias in the zero-cost budget predicate - PR #43512
Rate limiting​
- Add fail_closed_rate_limit_enforcement to reject requests with 503 while Redis rate limit counters are unreachable - PR #43251
MCP Gateway​
- Share compatibility-aware result conversion across tool surfaces - PR #43089
- Add Microsoft 365 (Graph) server to the MCP catalog - PR #43099
- Configure protocol versions and capability discovery - PR #43169
- Report reachability without stored credentials - PR #43240
- Align hub publication status and controls - PR #43241
- Adopt shared server resolution and caller authorization - PR #43263
- Scan and pin upstream tool descriptions - PR #43283
- Scope OpenAPI listings to the exact server prefix and drop upstream OAuth metadata when a server is saved - PR #43608
- Keep MCP permissions visible after key, team and MCP server saves - PR #43810
- Resolve team-granted toolsets for non-admin keys and dashboard sessions - PR #43908
Performance / Loadbalancing / Reliability improvements​
Auto Router and model routing​
- Parse the classifier verdict out of surrounding prose instead of falling to the default tier - PR #43215
- Opt in to prompt-cache cost routing - PR #43232
- Fetch cooldown state and usage counters in one Redis round trip - PR #43320
- Compare historical and new savings consistently - PR #43348
- Bind Claude Code background sessions to their auto-router - PR #43767
- Strip encrypted reasoning the pinned deployment cannot decrypt - PR #43781
- Carry per-request routing reads on context variables instead of public method kwargs - PR #43814
- Honour the cooldown read interval in the routing prefetch - PR #43815
- Show actual and baseline spend for historical savings - PR #44057
Caching, database and runtime​
- Add maximum_daily_tag_spend_retention_period cleanup setting - PR #39221
- Give SpendLogToolIndex its own share of the cleanup budget and log a per-run summary - PR #41768
- Propagate auth cache invalidation over Redis Cluster via a node-level pub/sub client - PR #43110
- Hold one spend counter batch across admission and across post-call accounting - PR #43369
- One request-scoped Redis pipeline for auth, spend, rate-limit and routing reads - PR #43407
- Run aresponses through the async wrapper so the cache is read once - PR #43769
- Refresh auth management objects through the request Redis pipeline - PR #43776
- One post-call Redis pipeline per backend for spend, rate-limit, routing and response-cache writes - PR #43779
- Build the SpendLogs indexes in the migration job instead of in migrations - PR #43948
- Write the response-cache SET to Redis at once instead of on the post-call batch - PR #43973
- Gzip buffered responses for clients that accept it - PR #44052
- Bound the lock waits of the partitioned SpendLogs index build - PR #44109
- Hand libpq a root cert, not Prisma's sslcert, when the migration job builds indexes - PR #44203
Dependency updates​
- Bump pyjwt, moment and brace-expansion to clear osv-scan - PR #43792
- Bump oauthlib to 4.0.0 to clear osv-scan - PR #43899
- Bump gitpython and tornado, extend diskcache osv ignore to Nov 1 - PR #43961
- Bump pypdf from 6.16.2 to 6.19.0 - PR #44033
Documentation Updates​
- Point readers to the security announcements mailing list signup - PR #43713
Tests, CI and Internal Changes​
These 98 PRs change tests, CI, contributor tooling, release packaging, or Rust runtime scaffolding that is not yet wired into a user-facing path. They do not change proxy or SDK behavior on their own
Tests (48)
- Typed per-test metadata for the e2e suite - PR #42044
- Record each e2e test's steps, starting with ProxyClient - PR #42393
- Resolve the blank-S3 gateway repo root from the litellm package location - PR #42911
- Pin async 5xx retry through the production AsyncHTTPHandler - PR #43080
- Move tests into active CI selection - PR #43235
- Pin the team-admin status-code matrix across every management route - PR #43249
- Pin server resolution and authorization behavior - PR #43261
- Classify every credential-bearing param for the canary suite - PR #43298
- Credential canary suite harness - PR #43300
- Sweep proxy logs, metrics, a Datadog intake and the Logs drawer for credential canaries - PR #43306
- Request-path credential canary slots D1-D4 - PR #43307
- Credential canary slots for MCP and pass-through credentials - PR #43308
- Stored-config credential canary slots - PR #43309
- Drain the logging worker after each logging callback test so no later test inherits its events - PR #43344
- Stop VCR recording and replaying a test's own localhost upstream - PR #43346
- Group /v1/messages contracts under tests/integration/messages_endpoint - PR #43352
- Remove substring guard test_default_api_base - PR #43355
- Native /v1/messages reasoning integration tests built on a captured Claude Code request - PR #43361
- Read management routes back from the control plane replicas - PR #43373
- Assert the usage the Gemma responses stream actually reports - PR #43422
- Enforce shared upstream error contract in wheel checks - PR #43520
- Pin org-admin status codes in the team-admin matrix - PR #43592
- Callback credential canary slots C1-C3 and D5 - PR #43630
- Refresh retired OpenAI tool-call models - PR #43676
- Accept regional aliases that inherit Converse routing - PR #43785
- Repair MCP Responses and budget fixtures - PR #43788
- Settle the shared logging worker before recording shadow callbacks - PR #43847
- Repair completion, SAIL, and spend-log fixtures - PR #43902
- Restore the AWS env after a failed live call in the auth tests - PR #43921
- Refresh qualified retired OpenAI fixtures - PR #43938
- Scope user_api_key_auth overrides in proxy_server tests - PR #43952
- Repair stale tests and move retired OpenAI text-completion fixtures - PR #43958
- Repair stale tests and flaky CI infrastructure - PR #43983
- Migrate DB and Redis backed proxy tests into tests/integration - PR #43996
- Move auth, hooks, policy_engine and client tests into tests/unit/proxy - PR #43998
- Move management_endpoints, management_helpers and guardrails tests into tests/unit/proxy - PR #44003
- Move utils, agent_endpoints and endpoint tests into tests/unit/proxy - PR #44006
- Inject the HIBP client and the MCP loop clock so two backend tests stop flaking - PR #44007
- Move proxy_server, _experimental and db tests into tests/unit/proxy - PR #44012
- Move middleware, spend_tracking, pass_through, common_utils and root proxy tests into tests/unit/proxy - PR #44015
- Delete the legacy proxy test tree and shard tests/unit/proxy by glob - PR #44018
- Bill Sail windows that synchronous calls can still use - PR #44058
- Run the db push timeout hint test without a database URL - PR #44073
- Keep 1ms-timeout deployments off the provider cache - PR #44082
- Move live-provider legacy tests into tests/e2e - PR #44120
- Move legacy proxy, router and Redis tests into tests/integration - PR #44128
- Assert a saved Straiker api_version v1 with an
sk_agt_key routes to v3 - PR #44153 - Repair stale and polluting tests red on scheduled main CI - PR #44229
CI (3)
Code quality and contributor tooling (19)
- Daily fresh tech debt cleanup, rolling PR (2026-09-25) - PR #43151
- Add shared server resolver without changing callers - PR #43262
- Replace Any with proven types in 5 files - PR #43304
- Add the backport-stable label only for a P0 regression - PR #43351
- Answer every team access check with TeamAccess.allows - PR #43364
- Consolidate hub publication predicate - PR #43394
- Clean up fresh tech debt from 2026-09-27 - PR #43538
- Replace Any with proven types in 8 files - PR #43551
- Drop UI, migration, and CODEOWNERS self owners - PR #43653
- Clean up fresh tech debt from 2026-09-28 - PR #43674
- Replace Any with proven types in 7 files - PR #43704
- Clean up fresh tech debt from 2026-09-29 - PR #43830
- Replace Any with proven types in 7 files - PR #43844
- Remove the LIT002 mutable-construction rule - PR #43971
- Clean up fresh tech debt from 2026-09-30 - PR #43993
- Point mcp_server test references at tests/unit/proxy - PR #44055
- Drop unused pytest-postgresql dev dependency - PR #44056
- Split the KeyActivityPanel condition chains to bring the lint budget back under its ceiling - PR #44114
- Remove banner comments, restating comments and dead in_loop_thread - PR #44161
Rust runtime internals (22)
- Add the openai_like chat config foundation - PR #43379
- Share anthropic types, request helpers, and streaming contracts across crates - PR #43426
- Expand gateway configuration parsing - PR #43460
- Share call lifecycle across route-owned inference - PR #43461
- Add the HTTP host driver - PR #43462
- Use shared execution in gateway inference - PR #43463
- Support the HTTP Responses API - PR #43464
- Connect Python inference bindings to shared routes - PR #43465
- Add structured route lifecycle tracing - PR #43466
- Separate gateway authentication and authorization - PR #43467
- Add virtual key storage contracts - PR #43468
- Add gateway UI login and sessions - PR #43469
- Add the MCP gateway - PR #43470
- Package the gateway container - PR #43471
- Add litellm-db and litellm-db-testing workspace scaffolding - PR #43504
- Remove delivery routing abstraction - PR #43514
- Centralize host execution and compose callbacks - PR #43515
- Select Rust caching through explicit cache objects - PR #43601
- Orchestrate Messages route execution - PR #43719
- Add shared llms wire type derives - PR #43730
- Centralize Python bridge execution wrappers - PR #43871
- Embed migration folders with a shared migrate! macro - PR #44104
Reverts of unreleased changes (2)
Release and packaging (4)
- Bump litellm-enterprise 0.1.71 -> 0.1.72, litellm-proxy-extras 0.4.102 -> 0.4.103, litellm 1.104.0 -> 1.105.0 - PR #43789
- Bump litellm-enterprise 0.1.72 -> 0.1.73, litellm-proxy-extras 0.4.103 -> 0.4.104 - PR #44126
- Bump litellm-proxy-extras 0.4.104 -> 0.4.105 - PR #44235
- Rebuild the Admin UI bundle on rc/1.105.0 - PR #44386
PR roll-up by ownership area​
Customer-facing PRs in rc.1: 238. Tests, CI and internal PRs: 98. Total: 336
- Models & Providers: 81
- Logging & Tracing: 36
- LLM API Endpoints: 27
- Performance: 26
- Auth & Management: 19
- Guardrails: 15
- UI: 13
- MCP: 10
- Spend / Budgets / Rate Limits: 10
- Docs: 1
New Contributors​
- @4refael made their first contribution in PR #41826
- @agustin18 made their first contribution in PR #43319
- @ankit373 made their first contribution in PR #43553
- @daqiangganjun made their first contribution in PR #38172
- @DeviaVir made their first contribution in PR #43558
- @fedaeho made their first contribution in PR #43512
- @Flexomatic81 made their first contribution in PR #43588
- @galovics made their first contribution in PR #42949
- @hsm207 made their first contribution in PR #43536
- @shoemoney made their first contribution in PR #38049
- @shrey-berri made their first contribution in PR #43221
- @stewartpark made their first contribution in PR #43147
- @YaseenBashaT made their first contribution in PR #43197
Full Changelog​
https://github.com/BerriAI/litellm/compare/v1.104.0-rc.1...v1.105.0-rc.1