LiteLLM Python SDK
The LiteLLM Python SDK is a Python library. It gives you one interface to call 100+ LLM providers, such as OpenAI, Anthropic, Vertex AI, and Bedrock, in the OpenAI format.
You import the SDK into your application code. You do not operate a server.
Install the SDK
uv add litellm
You can also use pip install litellm.
Make a request
Set the API key of your provider. Then call completion() with a model name in the format provider/model:
from litellm import completion
import os
os.environ["ANTHROPIC_API_KEY"] = "your-api-key"
response = completion(
model="anthropic/claude-sonnet-5",
messages=[{"role": "user", "content": "Hello, how are you?"}],
)
print(response.choices[0].message.content)
To use a different provider, change the model name and the API key. The code for the request and the response stays the same. For a full procedure, refer to the Quickstart.
What the SDK does
- One interface. Each provider uses the same
completion()function and the same parameters. - One output format. Each response has the OpenAI format, for all providers.
- Exception mapping. The SDK changes provider errors into the OpenAI exception types. Your code that catches OpenAI exceptions also catches the errors of each provider.
- Routing. The Router gives you retries, fallbacks, and load balancing across deployments.
- Callbacks. Send logs and costs to Langfuse, MLflow, Helicone, and other tools with one line of code.
SDK functions
Each function has an async version with the prefix a. For example, the async version of completion() is acompletion().
| Function | Use it to | Reference |
|---|---|---|
completion() | Send chat messages to a model | completion() |
responses() | Use the OpenAI Responses API format | responses() |
embedding() | Get vector embeddings for text | embedding() |
image_generation() | Make images from a text prompt | image_generation() |
transcription() | Change speech audio into text | transcription() |
speech() | Change text into speech audio | speech() |
rerank() | Put documents in order of relevance to a query | rerank() |
For all the endpoints that LiteLLM supports, refer to Supported Endpoints.
Frequent tasks
| Task | Page |
|---|---|
| Set API keys, API base URLs, and API versions | Set keys |
| Get responses as a stream | Streaming |
| Catch provider errors | Exception mapping |
| Send logs to observability tools | Callbacks |
| Calculate token usage and cost | Token usage |
| Add retries, fallbacks, and load balancing | Router |
| Cache responses | Caching |
| Run coding agents from Python | Agent harnesses |
SDK or AI Gateway
Use the SDK if you send requests to LLMs from one Python application. Use the AI Gateway if more than one application sends requests, or if virtual keys, spend tracking, or logs in one location are necessary.
The SDK can also send requests to an AI Gateway. Add the prefix litellm_proxy/ to the model name. Refer to LiteLLM Proxy as a provider.