Skip to main content

Custom HTTP Handler

Configure custom aiohttp sessions for better performance and control in LiteLLM completions.

Overview​

You can inject custom aiohttp.ClientSession instances into LiteLLM for:

  • Custom connection pooling and timeouts
  • Corporate proxy and SSL configurations
  • Performance optimization
  • Request monitoring
Scope

BaseLLMAIOHTTPHandler is only used by the aiohttp_openai/ provider (chat completions) and by Topaz image variations. Requests to other providers, including plain openai/, go through the httpx based clients and are not affected by this handler.

The instance that litellm.completion calls lives in the litellm.main module, so the replacement must be assigned to litellm.main.base_llm_aiohttp_handler. Setting litellm.base_llm_aiohttp_handler creates a new attribute on the litellm package that nothing reads, and the custom session is silently ignored. litellm.images.main binds its own reference to the handler at import time, so the Topaz image variation path is not changed by this assignment.

Basic Usage​

Default (No Changes Required)​

import litellm

# Works exactly as before
response = await litellm.acompletion(
model="gpt-5.6-luna",
messages=[{"role": "user", "content": "Hello!"}]
)

Custom Session​

import aiohttp
import litellm
import litellm.main
from litellm.llms.custom_httpx.aiohttp_handler import BaseLLMAIOHTTPHandler

# Create optimized session
session = aiohttp.ClientSession(
timeout=aiohttp.ClientTimeout(total=180),
connector=aiohttp.TCPConnector(limit=300, limit_per_host=75)
)

# Replace the handler that litellm.completion uses
litellm.main.base_llm_aiohttp_handler = BaseLLMAIOHTTPHandler(client_session=session)

# aiohttp_openai/ completions now use your session
response = await litellm.acompletion(model="aiohttp_openai/gpt-5.6-luna", messages=[...])

Common Patterns​

FastAPI Integration​

from contextlib import asynccontextmanager
from fastapi import FastAPI
import aiohttp
import litellm
import litellm.main
from litellm.llms.custom_httpx.aiohttp_handler import BaseLLMAIOHTTPHandler

@asynccontextmanager
async def lifespan(app: FastAPI):
# Startup
session = aiohttp.ClientSession(
timeout=aiohttp.ClientTimeout(total=180),
connector=aiohttp.TCPConnector(limit=300)
)
litellm.main.base_llm_aiohttp_handler = BaseLLMAIOHTTPHandler(
client_session=session
)
yield
# Shutdown
await session.close()

app = FastAPI(lifespan=lifespan)

@app.post("/chat")
async def chat(messages: list[dict]):
return await litellm.acompletion(model="aiohttp_openai/gpt-5.6-luna", messages=messages)

Corporate Proxy​

import ssl

# Custom SSL context
ssl_context = ssl.create_default_context()
ssl_context.load_cert_chain('cert.pem', 'key.pem')

# Proxy session
session = aiohttp.ClientSession(
connector=aiohttp.TCPConnector(ssl=ssl_context),
trust_env=True # Use environment proxy settings
)

litellm.main.base_llm_aiohttp_handler = BaseLLMAIOHTTPHandler(client_session=session)

High Performance​

# Optimized for high throughput
session = aiohttp.ClientSession(
timeout=aiohttp.ClientTimeout(total=300),
connector=aiohttp.TCPConnector(
limit=1000, # High connection limit
limit_per_host=200, # Per host limit
ttl_dns_cache=600, # DNS cache
keepalive_timeout=60, # Keep connections alive
enable_cleanup_closed=True
)
)

litellm.main.base_llm_aiohttp_handler = BaseLLMAIOHTTPHandler(client_session=session)

Constructor Options​

BaseLLMAIOHTTPHandler(
client_session=None, # Custom aiohttp.ClientSession
transport=None, # Advanced transport control
connector=None, # Custom aiohttp.BaseConnector
)

Resource Management​

  • User sessions: You manage the lifecycle (call await session.close())
  • Auto-created sessions: Automatically cleaned up by the handler
  • 100% backward compatible: Existing code works unchanged

Configuration Tips​

Development​

session = aiohttp.ClientSession(
timeout=aiohttp.ClientTimeout(total=60),
connector=aiohttp.TCPConnector(limit=50)
)

Production​

session = aiohttp.ClientSession(
timeout=aiohttp.ClientTimeout(total=300),
connector=aiohttp.TCPConnector(
limit=1000,
limit_per_host=200,
keepalive_timeout=60
)
)