LiteLLM is moving to Rust Read the latest updates.
v1.102.4 - Fix Double-Counted Spend on the Chat to Responses Bridge
Deploy this version
- Docker
- Pip
docker run \
-e LITELLM_MASTER_KEY=sk-<paste-a-long-random-key> \
-e DATABASE_URL=postgresql://<user>:<password>@<host>:5432/<dbname> \
-e STORE_MODEL_IN_DB=True \
-p 4000:4000 \
docker.litellm.ai/berriai/litellm:1.102.4
pip install litellm==1.102.4
This release is published as ghcr.io/berriai/litellm:v1.102.4. See the GitHub release and the full releases page
v1.102.4 is a patch release on top of v1.102.3. It fixes spend that was logged twice for large non-streaming chat requests. There are no new database migrations or breaking changes. The v1.102.4 tag points at 0de9a17
Spend logged once on the Chat to Responses bridge
Since v1.102.0, a large non-streaming /v1/chat/completions request to an openai/responses/* model was logged twice. The request showed two rows on the logs page, and key, team and daily spend went up by twice the real cost
- One request now queues exactly one success log, and the logged result is the provider's own response
- Streaming requests and the success handler itself are unchanged
What's Changed
- Log spend once for large non-streaming requests on the chat to Responses bridge - PR #44508, backported in PR #45215
Full Changelog
https://github.com/BerriAI/litellm/compare/v1.102.3...v1.102.4