Skip to main content
LiteLLM is moving to Rust Read the latest updates.

v1.102.4 - Fix Double-Counted Spend on the Chat to Responses Bridge

Deploy this version​

docker run \
-e LITELLM_MASTER_KEY=sk-<paste-a-long-random-key> \
-e DATABASE_URL=postgresql://<user>:<password>@<host>:5432/<dbname> \
-e STORE_MODEL_IN_DB=True \
-p 4000:4000 \
docker.litellm.ai/berriai/litellm:1.102.4

This release is published as ghcr.io/berriai/litellm:v1.102.4. See the GitHub release and the full releases page

v1.102.4 is a patch release on top of v1.102.3. It fixes spend that was logged twice for large non-streaming chat requests. There are no new database migrations or breaking changes. The v1.102.4 tag points at 0de9a17

Spend logged once on the Chat to Responses bridge​

Since v1.102.0, a large non-streaming /v1/chat/completions request to an openai/responses/* model was logged twice. The request showed two rows on the logs page, and key, team and daily spend went up by twice the real cost

  • One request now queues exactly one success log, and the logged result is the provider's own response
  • Streaming requests and the success handler itself are unchanged

What's Changed​

  • Log spend once for large non-streaming requests on the chat to Responses bridge - PR #44508, backported in PR #45215

Full Changelog​

https://github.com/BerriAI/litellm/compare/v1.102.3...v1.102.4

LiteLLM Enterprise
SSO/SAML, audit logs, spend tracking, multi-team management, and guardrails, built for production.
Learn more →