Skip to main content

26 posts tagged with "engineering"

View All Tags

How we cut LiteLLM's Redis round trips per request by 64%

Yassin Kortam
Senior SWE @ LiteLLM

LiteLLM's 22 Redis round trips merge into 8, a 64% reduction per request

This week, we cut the number of times a request to the LiteLLM proxy waits on Redis from 22 to 8.

A request from a key with a budget, in a team with a budget and TPM/RPM limits, against a model group with usage-based routing and a Redis response cache, made 22 Redis round trips: 12 before the model was called and 10 after. The same request, with the same checks and the same writes, now makes 5 before and 3 after.

JEV Classifier: 5.43x as Fast as Haiku, 96% Lower Cost

Moe Khalil
Product Engineer, LiteLLM

Last Updated: September 18, 2026

An Auto Router pays for classification before the selected model can answer. In our benchmark, TypeSafe JEV classified requests 5.43x as fast as Haiku, comparing median classifier latency: 126.81 ms versus 688.40 ms. Registry-priced classifier cost was 96.12% lower, rounded to 96% in the title

JEV matched our benchmark's expected tiers on 95.00% of calls, versus 73.75% for Haiku. That result depends on the prompts, tier definitions, instructions and context used here. The expected tiers were authored with the synthetic prompts, without independent review. This comparison does not establish general classification accuracy or the quality of the final answers

Auto-Router: Switching Tiers Without Encrypted Content Failures

Tin Lo
Founding AI Product Engineer, LiteLLM

Auto-Router can move a conversation between tiers as the request changes from simple work to a harder task. Responses API clients can make that switch difficult because a previous response may include encrypted reasoning content that only the deployment that created it can decrypt

LiteLLM now keeps the readable history and removes encrypted reasoning that the newly selected tier cannot verify. The request can continue to the selected model instead of failing with invalid_encrypted_content

This fix is included in the LiteLLM v1.102.x release line and landed in PR #40280

How Pfizer Improved LiteLLM Gateway Performance and Resiliency at Scale

Ishaan Jaffer
CTO, LiteLLM
Krrish Dholakia
CEO, LiteLLM
Yassin Kortam
Senior SWE @ LiteLLM
Pfizer AI Platform Engineering
Pfizer AI Platform Engineering
Pfizer AI Platform Engineering
Pfizer AI Platform Engineering
Pfizer AI Platform Engineering

LiteLLM x Pfizer

A joint debugging story with LiteLLM, and the release testing that comes next.

A LiteLLM version bump exposed a long-standing Redis configuration bug that cut Pfizer's gateway throughput by ~48%, with zero HTTP errors in the application logs. Here's how the team isolated it through configuration bisection, and the testing infrastructure both teams are building to catch this class of regression before it ships.

51% Cost Savings Reported From a Live Production Deployment

Tin Lo
Founding AI Product Engineer, LiteLLM

You can expect roughly 40% cost reductions from day one with the Auto Router, and more as the tier maps are tuned. One of our production users shared their statistics to show what that looks like at scale.

They rolled out the Auto-Router to 450+ users across dev, staging, and prod instances and saved $12,249 over 270k+ requests.

Production traffic recorded 51% cost-savings with the Auto-Router

Benchmarking the LiteLLM Rust AI Gateway: Overhead, Memory, and Cost

Ishaan Jaffer
CTO, LiteLLM

Last Updated: July 2026

We are launching an early beta of the LiteLLM AI Gateway in Rust, and we built AIGatewayBench to measure it against Portkey, Bifrost, and the current LiteLLM Python proxy. Across all four, the LiteLLM Rust gateway has the lowest p99 added latency and the smallest memory footprint by a wide margin: roughly 7x lower overhead and 9x less memory than the next-closest gateway (Bifrost), the lowest cost footprint, and the fastest whole-session times for coding agents. It holds its own on raw sustained throughput and pulls decisively ahead on overhead, memory, and cost.

Migrating LiteLLM to Rust - Building the Fastest and Litest AI Gateway

Ishaan Jaffer
CTO, LiteLLM

Last Updated: June 2026

Over the past year, we have heard the same thing from our users and our community: they want the fastest, most lightweight AI gateway they can run. We have heard you. We are addressing it by moving LiteLLM to Rust, and committing to sub-1ms overhead with a sub-100MB memory binary you can deploy. By the end of this migration, you will get a pure Rust server that can serve 100% of your AI traffic, with every hot path operation, including auth and rate limiting, running in Rust.

Want to help us build it?

We are opening an early beta and want to work directly with teams who care about a fast, lightweight gateway. If that is you, sign up here and we will get you testing the Rust gateway in your own stack, with a direct line to our team.

The reason it matters: under real load, CPU and memory climb with concurrency, and pods get OOM-killed at the worst time. Today the LiteLLM Python proxy peaks around 359MB of memory under load, and that cost multiplies across every pod, region, and retry you run.

We are already seeing the payoff in benchmarks. The Rust gateway serves about 15x the throughput (453 to 6,782 requests per second) on about 11x less memory (359MB to 32MB), and cuts per-request overhead from about 7.5ms on the Python path to about 0.05ms, well under the 1ms we commit to.

What you get​

You deploy a single Rust binary. It uses about 65MB of memory, gateway overhead stays under 1ms, and nothing in your setup changes: same config.yaml, same database, same client API, same providers. You keep LiteLLM's coverage of 100+ LLM providers behind one OpenAI-compatible API, with /chat/completions, /messages, /responses, and every other LLM endpoint LiteLLM supports today, now as the fastest and most lightweight LLM gateway you can self-host.

This is not a v2 and not a rewrite. There is no new major version to migrate to and nothing for you to change. The runtime under the hot path gets faster and lighter while your config stays exactly where it is.

We ship this the careful way. Each route moves to Rust only after it passes our full parity and end-to-end test suite, and it runs in production before the next route starts. Stability is the priority, and we target zero regressions on every release.

How we built a background agent to cover 30% of our backlog

Krrish Dholakia
CEO, LiteLLM
Ishaan Jaffer
CTO, LiteLLM
LiteLLM Agent Platform: agent.litellm.ai
info

The platform we built is open source. Check out litellm-agent-platform. The swappable harness layer is lite-harness.

Building the same thing inside your company?

Our goal was to 10x the productivity of our company with agents.

Three weeks ago we began building an agent that could own 30% of our engineering tickets. Here's what we've learnt so far.

Making the AI Gateway Resilient to Redis Failures

Ishaan Jaffer
CTO, LiteLLM

Last Updated: April 2026

Enterprise AI Gateway deployments put Redis in the hot path for nearly every request: rate limiting, cache lookups, spend tracking. When Redis is healthy, the latency contribution is single-digit milliseconds, invisible to end users. When it degrades, a production AI Gateway needs to stay up regardless.

Running LiteLLM at scale across 100+ pods means designing for failure modes before they appear. The easy case is Redis going fully down: fail fast, fall through to the database, continue serving requests. The hard case, the one that takes down gateways, is a slow Redis: still accepting connections, still responding, but timing out after 20-30 seconds per operation.