Skip to main content

Benchmarks

Benchmarks for LiteLLM Gateway (Proxy Server) tested against a fake OpenAI endpoint.

LiteLLM Gateway has 8ms P95 latency at 1k RPS (See benchmarks here)

High-throughput profile: 3,000 RPS with 50K to 100K-token prompts​

Large prompts create a different gateway workload than short chat requests. Token counting, budget checks, spend tracking, and metrics collection all happen before or after the model-provider call and can become bottlenecks at high request volume.

This benchmark compares the high-throughput deployment profile with v1.101.0. The profile combines Rust token counting, shared database connections, isolated metrics and spend processing, and traffic-based autoscaling.

Nightly benchmark

The high-throughput profile is still in development and is available in nightly builds. These results used the earliest available version of the complete profile.

Results​

CategoryMetricHigh-throughput profilev1.101.0Change
DeploymentGateway pods331324x fewer
Workers per pod41
Total workers132132same
ThroughputRequests/sec3.00K0.19K16x
Tokens/sec224.61M6.92M32x
Projected tokens/30 days582.20T17.94T32x
ReliabilityHTTP 200 rate (Locust)100.00%92.07%
Request latencyp5030.581 ms6.950 s227x
p9554.029 ms27.451 s508x
p9991.645 ms29.826 s325x
Time to first tokenp5031.667 ms9.400 s297x

The profile reached the full 3,000 RPS target with 100 percent client-visible success. The baseline settled near 190 RPS and returned a successful response for 92.07 percent of requests.

Test setup​

Test dimensionConfiguration
Load generatorDistributed Locust with one master and 30 workers
Traffic3,000 simulated users at one request per second each
Request mix50K, 75K, and 100K-token prompts in equal shares
Streaming50 percent of requests
Endpoint/v1/chat/completions with max_tokens: 16
AuthenticationVirtual key with a budget, so admission token counting and budget reservation ran
ModelIn-process mock model with response caching disabled
Network pathPublic AWS Application Load Balancer
Client timeout60 seconds
Run durationHigh-throughput profile: 24m 22s. Baseline: 5m 7s.

The mock model removes provider cost and provider latency while keeping the gateway request path active. The test still includes authentication, budgets, token counting, spend tracking, and metrics.

Both deployments ran 132 total gateway workers and requested 528 GiB of memory. The high-throughput profile used 33 pods with four workers per pod and requested 132 vCPU. The baseline used 132 pods with one worker per pod and requested 264 vCPU.

What made the difference​

Each change below was measured separately before the complete profile was tested.

ChangeCustomer impactMeasured effect
Rust admission token countingReduces CPU spent counting large prompts before dispatch.50K / 75K / 100K counts fell from 46 / 53 / 100 ms to 4.9 / 6.8 / 10.2 ms.
PgBouncer per podPrevents database connections from multiplying with every worker.Postgres held 86 to 175 connections across 11 to 29 pods, with no waiting PgBouncer clients.
Spend collector sidecarKeeps spend processing away from inference workers.At 700 RPS, p99 fell from 1.8 s to 830 ms. Total compute stayed roughly the same.
Metrics sidecarKeeps Prometheus scrapes away from inference workers.The sidecar used about 2 millicores per pod at 700 RPS.
Higher CPU burst limitPrevents all workers in a pod from being throttled together.At 700 RPS, p99 fell from 830 ms to 670 ms.
Gateway keep-aliveKeeps load-balancer connections valid during scaling.ALB-generated 502 responses fell from 15 to zero in the 200 RPS test.
RPS and TPS autoscalingReacts to traffic before CPU becomes saturated.A new replica was added about 48 seconds after a 200-user load step.
Admission token-count reuseAvoids counting the same large streaming prompt twice in mock tests.Streaming mock requests finished within about 30 ms of non-streaming requests.

How to read the metrics​

  • Requests per second, tokens per second, projected tokens, and request latency come from the gateway's Prometheus metrics.
  • Time to first token comes from Locust and measures the time from sending the request to receiving the first streaming event. It includes request upload, the load balancer, and gateway admission work.
  • The HTTP 200 rate comes from Locust because it includes failures that never reached the gateway.

The v1.101.0 run had 5,118 client-visible failures: 4,546 client timeouts or dropped connections, 457 HTTP 504 responses, and 115 HTTP 502 responses. The gateway did not receive these requests, so its own success metric showed 100 percent while Locust showed 92.07 percent.

Use the POST rows when reading Locust throughput. Each streaming request also creates a TTFT row, so the Locust Aggregated row counts more entries than real requests when streaming is enabled.

Benchmark scope​

This is a before-and-after comparison of the complete profile, not a single-variable test. The deployments used different pod shapes and ran for different lengths of time. The individual effects in the table above come from separate A/B tests at 200 to 1,000 RPS.

The in-process mock model excludes provider latency. These results measure gateway capacity for this specific traffic shape and should not be treated as universal production sizing guidance. Measure a representative workload before choosing worker counts, pod resources, and HPA targets.

The sections below use short request bodies against a fake OpenAI endpoint on 4 CPU / 8 GB machines. They are not directly comparable with this large-prompt benchmark.

Machine Spec used for testing​

Each machine deploying LiteLLM had the following specs:

  • 4 CPU
  • 8GB RAM

Configuration​

  • Database: PostgreSQL. See Database Sizing for how to size yours
  • Redis: Not used. Recommended in production; see Redis Sizing
  • Load generator: Locust, 1000 users, each with 0.5s to 1s of think time between requests. See Locust Settings before comparing these numbers against your own run.

2 Instance LiteLLM Proxy​

In these tests the baseline latency characteristics are measured against a fake-openai-endpoint.

Performance Metrics​

TypeNameMedian (ms)95%ile (ms)99%ile (ms)Average (ms)Current RPS
POST/chat/completions2006301200262.461035.7
CustomLiteLLM Overhead Duration (ms)12294314.741035.7
Aggregated100430930138.62071.4

4 Instances​

TypeNameMedian (ms)95%ile (ms)99%ile (ms)Average (ms)Current RPS
POST/chat/completions100150240111.731170
CustomLiteLLM Overhead Duration (ms)28133.321170
Aggregated7713018057.532340

Key Findings​

  • Doubling from 2 to 4 LiteLLM instances halves median latency: 200 ms → 100 ms.
  • High-percentile latencies drop significantly: P95 630 ms → 150 ms, P99 1,200 ms → 240 ms.
  • Setting workers equal to CPU count gives optimal performance.

Setting Up Benchmarking with Network Mock​

The fastest way to benchmark proxy overhead is using network_mock mode. This intercepts outbound requests at the httpx transport layer and returns canned responses, no need for setting up a mock provider.

1. Create a proxy config:

model_list:
- model_name: db-openai-endpoint
litellm_params:
model: openai/gpt-5.6-terra
api_key: "sk-fake-key"
api_base: "https://api.openai.com"

litellm_settings:
network_mock: true
callbacks: []
num_retries: 0
request_timeout: 30

general_settings:
master_key: "sk-<your-litellm-master-key>"

2. Start the proxy:

litellm --config benchmark_config.yaml --port 4000 --num_workers 8

3. Run the benchmark script:

python scripts/benchmark_mock.py --requests 2000 --max-concurrent 200 --runs 3

Get the benchmarking script here

This measures pure proxy overhead on the hot path without any network latency to a real or fake provider.

Setting Up a Fake OpenAI Endpoint​

For load testing and benchmarking, you can use a fake OpenAI proxy server. LiteLLM provides:

  1. Hosted endpoint: Use our free hosted fake endpoint at https://exampleopenaiendpoint-production.up.railway.app/
  2. Self-hosted: Set up your own fake OpenAI proxy server using github.com/BerriAI/example_openai_endpoint

Use this config for testing:

model_list:
- model_name: "fake-openai-endpoint"
litellm_params:
model: openai/any
api_base: https://exampleopenaiendpoint-production.up.railway.app/ # or your self-hosted endpoint
api_key: "test"

/realtime API Benchmarks​

End-to-end latency benchmarks for the /realtime endpoint tested against a fake realtime endpoint.

Performance Metrics​

MetricValue
Median latency59 ms
p95 latency67 ms
p99 latency99 ms
Average latency63 ms
RPS1,207

Test Setup​

CategorySpecification
Load TestingLocust: 1,000 users with 0.5s to 1s think time, 500 ramp-up
System4 vCPUs, 8 GB RAM, 4 workers, 4 instances
DatabasePostgreSQL (Redis unused)

Infrastructure Recommendations​

The runs above used a single PostgreSQL instance and no Redis, which is a benchmark configuration rather than a production one. For instance sizes at each request rate, the connection math that decides whether a deployment survives a scale-out, and concrete managed-service picks on AWS, Azure, and GCP, see Database Sizing and Redis Sizing. For the gateway-side configuration that goes with it, see Production Best Practices.

Locust Settings​

  • 1000 Users
  • 500 user Ramp Up
  • wait_time = between(0.5, 1), so every user sleeps 0.5s to 1s between requests

Why the think time matters when you reproduce these numbers​

A Locust user spends its time either waiting on a response or sleeping. With a 0.75s mean think time and ~110ms responses, each of the 1000 users completes a request about every 0.86s, so the run offers ~1160 RPS and holds roughly 130 requests in flight at any instant. That in-flight depth, not the user count, is what the latency columns above describe.

A closed-loop client with no think time is measuring something else. 1000 concurrent workers that send the next request the moment the previous one returns hold 1000 requests in flight, about 8x the queue depth of these runs. Once a gateway is saturated its throughput is fixed, and by Little's Law the latency each client observes is just requests in flight / throughput. So the same deployment, at the same RPS, reports roughly 8x the latency purely because the client queued 8x as much work into it. Latency and concurrency are not independent, and neither number means anything without the other.

To compare against the tables above, either keep the 0.5s to 1s think time, or hold your client's in-flight request count near 130 and report it alongside the latency. It is also worth reporting RPS first: if your run shows higher RPS and higher latency than these tables, your gateway is faster than this benchmark and your client is simply queueing deeper.

How to measure LiteLLM Overhead​

All responses from litellm will include the x-litellm-overhead-duration-ms header, this is the latency overhead in milliseconds added by LiteLLM Proxy.

If you want to measure this on locust you can use the following code:

Locust Code for measuring LiteLLM Overhead
import os
import uuid
from locust import HttpUser, task, between, events

# Custom metric to track LiteLLM overhead duration
overhead_durations = []

@events.request.add_listener
def on_request(request_type, name, response_time, response_length, response, context, exception, start_time, url, **kwargs):
if response and hasattr(response, 'headers'):
overhead_duration = response.headers.get('x-litellm-overhead-duration-ms')
if overhead_duration:
try:
duration_ms = float(overhead_duration)
overhead_durations.append(duration_ms)
# Report as custom metric
events.request.fire(
request_type="Custom",
name="LiteLLM Overhead Duration (ms)",
response_time=duration_ms,
response_length=0,
)
except (ValueError, TypeError):
pass

class MyUser(HttpUser):
wait_time = between(0.5, 1) # Random wait time between requests

def on_start(self):
self.api_key = os.getenv('API_KEY', 'sk-<your-litellm-api-key>')
self.client.headers.update({'Authorization': f'Bearer {self.api_key}'})

@task
def litellm_completion(self):
# no cache hits with this
payload = {
"model": "db-openai-endpoint",
"messages": [{"role": "user", "content": f"{uuid.uuid4()} This is a test there will be no cache hits and we'll fill up the context" * 150}],
"user": "my-new-end-user-1"
}
response = self.client.post("chat/completions", json=payload)

if response.status_code != 200:
# log the errors in error.txt
with open("error.txt", "a") as error_log:
error_log.write(response.text + "\n")

LiteLLM vs Portkey Performance Comparison​

Test Configuration: 4 CPUs, 8 GB RAM per instance | Load: 1k concurrent users, 500 ramp-up Versions: Portkey v1.14.0 | LiteLLM v1.79.1-stable
Test Duration: 5 minutes

Multi-Instance (4×) Performance​

MetricPortkey (no DB)LiteLLM (with DB)Comment
Total Requests293,796312,405LiteLLM higher
Failed Requests00Same
Median Latency100 ms100 msSame
p95 Latency230 ms150 msLiteLLM lower
p99 Latency500 ms240 msLiteLLM lower
Average Latency123 ms111 msLiteLLM lower
Current RPS1,170.91,170Same

Lower is better for latency metrics; higher is better for requests and RPS.

Technical Insights​

Portkey

Pros

  • Low memory footprint
  • Stable latency with minimal spikes

Cons

  • CPU utilization capped around ~40%, indicating underutilization of available compute resources
  • Experienced three I/O timeout outages

LiteLLM

Pros

  • Fully uses available CPU capacity
  • Strong connection handling and low latency after initial warm-up spikes

Cons

  • High memory usage during initialization and per request

Logging Callbacks​

GCS Bucket Logging​

Using GCS Bucket has no impact on latency, RPS compared to Basic Litellm Proxy

MetricBasic Litellm ProxyLiteLLM Proxy with GCS Bucket Logging
RPS1133.21137.3
Median Latency (ms)140138

LangSmith logging​

Using LangSmith has no impact on latency, RPS compared to Basic Litellm Proxy

MetricBasic Litellm ProxyLiteLLM Proxy with LangSmith
RPS1133.21135
Median Latency (ms)140132