
LiteLLM was counting every token in a long conversation to answer a yes-or-no routing question.
Removing that unnecessary work took median time to first byte from 553 ms to 35 ms in our local 440k-token benchmark: 94% lower.
Counting 440,000 tokens to check for 1,024
LiteLLM's optional prompt_caching routing check helps keep requests on deployments that can reuse their cached prompt. To check whether a prompt met a deployment's minimum size, often 1,024 tokens, it counted the entire conversation.
For a coding session with hundreds of turns and tool results, that could mean hundreds of milliseconds of Python work before the model call.
We skip the check for one healthy deployment. There is no routing choice to make, so we now skip the eligibility count, prefix hash, and cache-pin lookup entirely. This is the path measured below.
With multiple deployments, counting stops once the answer is known. The eligibility check stops after the first message that brings the count to the required minimum. Cache affinity and prefix hashing remain in place. Other uses of token counting, including usage accounting, are unchanged.
553 ms → 35 ms
The benchmark used the Python request path, a 439,945-token conversation with 334 turns and 18 tools, one healthy deployment, Redis response caching, and the prompt_caching check enabled. The test provider replied immediately, isolating request overhead from model generation time.
Each result is the median of three requests after one warmup, on the same local machine. /v1/responses already bypassed this check and stayed roughly flat. Requests without the optional check are unaffected. Full benchmark samples and setup.
See the changes: prompt-cache routing optimization.
