Restrict Model Access
Restrict models by Virtual Key
Set allowed models for a key using the models param
The models list both hides a model from GET /v1/models and blocks calls to it. To hide a model from the listing endpoints without blocking calls to it, set model_info.discoverable: false on the model instead (Hide a model from /v1/models)
curl 'http://0.0.0.0:4000/key/generate' \
--header 'Authorization: Bearer <your-master-key>' \
--header 'Content-Type: application/json' \
--data-raw '{"models": ["gpt-5.6-luna", "gpt-5.6-terra"]}'
This key can only make requests to models that are gpt-5.6-luna or gpt-5.6-terra
Verify this is set correctly by
- Allowed Access
- Disallowed Access
curl -i http://localhost:4000/v1/chat/completions \
-H "Content-Type: application/json" \
-H "Authorization: Bearer $LITELLM_API_KEY" \
-d '{
"model": "gpt-5.6-terra",
"messages": [
{"role": "user", "content": "Hello"}
]
}'
Expect this to fail since claude-sonnet-5 is not in the models for the key generated
curl -i http://localhost:4000/v1/chat/completions \
-H "Content-Type: application/json" \
-H "Authorization: Bearer $LITELLM_API_KEY" \
-d '{
"model": "claude-sonnet-5",
"messages": [
{"role": "user", "content": "Hello"}
]
}'
API Reference
Restrict models by team_id
litellm-dev can only access azure-gpt-3.5
1. Create a team via /team/new
curl --location 'http://localhost:4000/team/new' \
--header 'Authorization: Bearer <your-master-key>' \
--header 'Content-Type: application/json' \
--data-raw '{
"team_alias": "litellm-dev",
"models": ["azure-gpt-3.5"]
}'
# returns {...,"team_id": "my-unique-id"}
2. Create a key for team
curl --location 'http://localhost:4000/key/generate' \
--header "Authorization: Bearer $LITELLM_API_KEY" \
--header 'Content-Type: application/json' \
--data-raw '{"team_id": "my-unique-id"}'
3. Test it
curl --location 'http://0.0.0.0:4000/chat/completions' \
--header 'Content-Type: application/json' \
--header 'Authorization: Bearer sk-qo992IjKOC2CHKZGRoJIGA' \
--data '{
"model": "BEDROCK_GROUP",
"messages": [
{
"role": "user",
"content": "hi"
}
]
}'
{"error":{"message":"Invalid model for team litellm-dev: BEDROCK_GROUP. Valid models for team are: ['azure-gpt-3.5']\n\n\nTraceback (most recent call last):\n File \"/Users/ishaanjaffer/Github/litellm/litellm/proxy/proxy_server.py\", line 2298, in chat_completion\n _is_valid_team_configs(\n File \"/Users/ishaanjaffer/Github/litellm/litellm/proxy/utils.py\", line 1296, in _is_valid_team_configs\n raise Exception(\nException: Invalid model for team litellm-dev: BEDROCK_GROUP. Valid models for team are: ['azure-gpt-3.5']\n\n","type":"None","param":"None","code":500}}%
API Reference
View Available Fallback Models
Use the /v1/models endpoint to discover available fallback models for a given model. This helps you understand which backup models are available when your primary model is unavailable or restricted.
The include_metadata parameter serves as an extension point for exposing additional model metadata in the future. While currently focused on fallback models, this approach will be expanded to include other model metadata such as pricing information, capabilities, rate limits, and more.
Basic Usage
Get all available models:
curl -X GET 'http://localhost:4000/v1/models' \
-H 'Authorization: Bearer <your-api-key>'
Get Fallback Models with Metadata
Include metadata to see fallback model information:
curl -X GET 'http://localhost:4000/v1/models?include_metadata=true' \
-H 'Authorization: Bearer <your-api-key>'
Get Specific Fallback Types
You can specify the type of fallbacks you want to see:
- General Fallbacks
- Context Window Fallbacks
- Content Policy Fallbacks
curl -X GET 'http://localhost:4000/v1/models?include_metadata=true&fallback_type=general' \
-H 'Authorization: Bearer <your-api-key>'
General fallbacks are alternative models that can handle the same types of requests.
curl -X GET 'http://localhost:4000/v1/models?include_metadata=true&fallback_type=context_window' \
-H 'Authorization: Bearer <your-api-key>'
Context window fallbacks are models with larger context windows that can handle requests when the primary model's context limit is exceeded.
curl -X GET 'http://localhost:4000/v1/models?include_metadata=true&fallback_type=content_policy' \
-H 'Authorization: Bearer <your-api-key>'
Content policy fallbacks are models that can handle requests when the primary model rejects content due to safety policies.
Example Response
When include_metadata=true is specified, each model carries a metadata.fallbacks list for a single fallback type, the one named by fallback_type (general when omitted). To see all three types, send one request per fallback_type:
{
"data": [
{
"id": "gpt-5.6-terra",
"object": "model",
"created": 1677610602,
"owned_by": "openai",
"metadata": {
"fallbacks": ["gpt-5.6-luna", "claude-sonnet-5"]
}
}
]
}
Use Cases
- High Availability: Identify backup models to ensure service continuity
- Cost Optimization: Find cheaper alternatives when primary models are expensive
- Content Filtering: Discover models with different content policies
- Context Length: Find models that can handle larger inputs
- Load Balancing: Distribute requests across multiple compatible models
API Parameters
| Parameter | Type | Description |
|---|---|---|
include_metadata | boolean | Include additional model metadata including fallbacks |
fallback_type | string | Which fallbacks to return in metadata.fallbacks: general (default), context_window, or content_policy. Any other value returns a 400 |
Reserve a deployment for a team during a time window
Set model_info.access_windows on a deployment to reserve it for specific teams during a daily local-time window. While a window is active the router only hands that deployment to requests whose key belongs to one of the listed teams; requests from other teams, and requests from keys with no team (including the master key), are routed to other deployments in the same model group or rejected with a 400 if every candidate is reserved. Outside the window routing is unchanged. The deployment stays listed in /v1/models and /model/info at all times.
model_list:
- model_name: gpt-4o-ptu
litellm_params:
model: azure/gpt-4o-ptu
api_base: os.environ/AZURE_PTU_BASE
api_key: os.environ/AZURE_PTU_KEY
model_info:
access_windows:
- start: "22:00"
end: "06:00"
timezone: "America/New_York"
team_ids: ["team-nightly-batch"]
start and end are HH:MM wall-clock times in the given IANA timezone (daylight saving is applied automatically). start is inclusive and end is exclusive; a start later than end means the window crosses midnight. A deployment can list several windows; it is reserved whenever any of them is active. The proxy refuses to start when a window has an invalid time, an unknown timezone, an empty team_ids, or equal start and end.
A rejected request looks like this:
{"error":{"message":"litellm.BadRequestError: Deployment gpt-4o-ptu is reserved for another team until 06:00 America/New_York","type":"invalid_request_error","param":null,"code":"400"}}
Advanced: Model Access Groups
For advanced use cases, use Model Access Groups to dynamically group multiple models and manage access without restarting the proxy.