Skip to main content

Getting Started

Who calls
Developer
Coding agent
Your app
LiteLLM
LLM APIs
MCP tools
A2A agents
Quick Start
curl -fsSL https://raw.githubusercontent.com/BerriAI/litellm/main/scripts/quickstart.sh | sh

LiteLLM is an open-source library that gives you a single, unified interface to call 100+ LLMs (OpenAI, Anthropic, Vertex AI, Bedrock, and more) using the OpenAI format.

  • Call any provider using the same completion() interface, with no API to re-learn for each one
  • Consistent output format regardless of which provider or model you use
  • Built-in retry / fallback logic across multiple deployments via the Router
  • Self-hosted LLM Gateway (Proxy) with virtual keys, cost tracking, and an admin UI

PyPI GitHub Stars


Installation​

uv add litellm

A leaner SDK installation​

For applications that use the Python SDK directly, litellm-core provides the shared SDK with fewer mandatory dependencies and without the bundled dashboard or gateway CLI entry points. Python imports remain unchanged: continue using litellm

Install litellm-core in a fresh environment instead of installing litellm. Add AWS and Python Hugging Face tokenizer packages when your application needs them. The two distributions cannot be installed together because they share the same Python files

See LiteLLM Core for installation, package selection, and optional dependencies. Existing litellm installations retain their current dependency defaults

To deploy the full AI Gateway (Proxy) with the Admin UI, follow the Quickstart; it runs as a container and needs no Python setup. To run it from the CLI instead, see the Gateway Quickstart.


Quick Start​

Make your first LLM call using the provider of your choice:

from litellm import completion
import os

os.environ["OPENAI_API_KEY"] = "your-api-key"

response = completion(
model="openai/gpt-5.6-terra",
messages=[{"role": "user", "content": "Hello, how are you?"}]
)
print(response.choices[0].message.content)

Every response follows the OpenAI Chat Completions format, regardless of provider. ✅

Response Format​

Non-streaming responses return a ModelResponse object:

{
"id": "chatcmpl-abc123",
"object": "chat.completion",
"created": 1677858242,
"model": "gpt-5.6-terra",
"choices": [
{
"index": 0,
"message": {
"role": "assistant",
"content": "Hello! I'm doing well, thanks for asking."
},
"finish_reason": "stop"
}
],
"usage": {
"prompt_tokens": 13,
"completion_tokens": 12,
"total_tokens": 25
}
}

Streaming responses (stream=True) yield ModelResponseStream chunks:

{
"id": "chatcmpl-abc123",
"object": "chat.completion.chunk",
"created": 1677858242,
"model": "gpt-5.6-terra",
"choices": [
{
"index": 0,
"delta": {
"role": "assistant",
"content": "Hello"
},
"finish_reason": null
}
]
}

📖 Full output format reference →

Open in Colab
Open In Colab

New to LiteLLM?​

Want to get started fast? Head to Tutorials for step-by-step walkthroughs of AI coding tools, agent SDKs, proxy setup, and more.

Need to understand a specific feature? Check Guides for streaming, function calling, prompt caching, and other how-tos.


Choose Your Path​


LiteLLM Python SDK​

Streaming​

Add stream=True to receive chunks as they are generated:

from litellm import completion
import os

os.environ["OPENAI_API_KEY"] = "your-api-key"

for chunk in completion(
model="openai/gpt-5.6-terra",
messages=[{"role": "user", "content": "Write a short poem"}],
stream=True,
):
print(chunk.choices[0].delta.content or "", end="")

Exception Handling​

LiteLLM maps every provider's errors to the OpenAI exception types, so your existing error handling keeps working:

import litellm

try:
litellm.completion(
model="anthropic/claude-sonnet-5",
messages=[{"role": "user", "content": "Hey!"}]
)
except litellm.AuthenticationError as e:
print(f"Bad API key: {e}")
except litellm.RateLimitError as e:
print(f"Rate limited: {e}")
except litellm.APIError as e:
print(f"API error: {e}")

Logging & Observability​

Send input/output to Langfuse, MLflow, Helicone, Lunary, and more with a single line:

import litellm

litellm.success_callback = ["langfuse", "mlflow", "helicone"]

response = litellm.completion(
model="gpt-5.6-terra",
messages=[{"role": "user", "content": "Hi!"}]
)

📖 See all observability integrations →

Track Costs & Usage​

Use a callback to capture cost per response:

import litellm

def track_cost(kwargs, completion_response, start_time, end_time):
print("Cost:", kwargs.get("response_cost", 0))

litellm.success_callback = [track_cost]

litellm.completion(
model="gpt-5.6-terra",
messages=[{"role": "user", "content": "Hello!"}],
stream=True
)

📖 Custom callback docs →


LiteLLM Proxy Server (LLM Gateway)​

The proxy is a self-hosted OpenAI-compatible gateway. Any client that works with OpenAI works with the proxy, with no code changes.

LiteLLM Proxy Dashboard

Step 1: Start the proxy​

litellm --model huggingface/bigcode/starcoder
# Proxy running on http://0.0.0.0:4000

Step 2: Call it with the OpenAI client​

import openai

client = openai.OpenAI(api_key="anything", base_url="http://0.0.0.0:4000")

response = client.chat.completions.create(
model="gpt-5.6-luna",
messages=[{"role": "user", "content": "Write a short poem"}]
)
print(response.choices[0].message.content)

👉 Full proxy quickstart →

Debugging tool

Use /utils/transform_request to inspect exactly what LiteLLM sends to any provider. It helps when debugging prompt formatting, header issues, and provider-specific parameters.

🔗 Interactive API explorer (Swagger) →


Agent & MCP Gateway​

LiteLLM is a unified gateway for LLMs, agents, and MCP, so you don't need a separate agent or MCP gateway. One endpoint for 100+ models, A2A agents, and MCP tools.


What to Explore Next​