---
title: "Prompt Caching"
url: "/docs/completion/prompt_caching"
canonical_url: "https://docs.litellm.ai/docs/completion/prompt_caching"
type: "docs"
last_updated: "2026-10-07"
related:
  - "/docs/completion/message_trimming"
  - "/docs/completion/prompt_formatting"
---
# Prompt Caching

> Index of all LiteLLM docs: https://docs.litellm.ai/llms.txt


Supported Providers:
- OpenAI (`openai/`)
- Anthropic API (`anthropic/`)
- Google AI Studio (`gemini/`)
- Vertex AI (`vertex_ai/`, `vertex_ai_beta/`)
- Bedrock (`bedrock/`, `bedrock/invoke/`, `bedrock/converse`) ([All models bedrock supports prompt caching on](https://docs.aws.amazon.com/bedrock/latest/userguide/prompt-caching.html))
- Deepseek API (`deepseek/`)
- xAI (`xai/`)

:::warning[Minimum token requirements]
Prompt caching is silently skipped when the input is below the provider's minimum, and **no error is returned**. Always verify caching occurred by checking `cache_creation_input_tokens` in the response.

| Provider | Minimum input tokens |
|---|---|
| OpenAI | 1,024 |
| Anthropic (Claude Opus 5, Fable 5, Mythos 5) | 512 |
| Anthropic (Claude Sonnet 5, Opus 4.8, Sonnet 4.x, Opus 4, 4.1, Claude 3.x) | 1,024 |
| Anthropic (Claude Haiku 4.5, Opus 4.5, 4.6) | 4,096 |
| Bedrock (Claude Opus 5) | 512 |
| Bedrock (Claude Sonnet 5, Opus 4.8, Sonnet 4.x, Claude 3.5, 3.7) | 1,024 |
| Bedrock (Claude Haiku 4.5, Opus 4.5, 4.6, 4.7) | 4,096 |
| Google Gemini | 1,024 |
:::

For the supported providers, LiteLLM follows the OpenAI prompt caching usage object format:

```bash
"usage": {
  "prompt_tokens": 2006,
  "completion_tokens": 300,
  "total_tokens": 2306,
  "prompt_tokens_details": {
    "cached_tokens": 1920
  },
  "completion_tokens_details": {
    "reasoning_tokens": 0
  }
  # ANTHROPIC_ONLY #
  "cache_creation_input_tokens": 0
}
```

- `prompt_tokens`: These are all prompt tokens including cache-miss and cache-hit input tokens.
- `completion_tokens`: These are the output tokens generated by the model.
- `total_tokens`: Sum of prompt_tokens + completion_tokens.
- `prompt_tokens_details`: Object containing cached_tokens.
    - `cached_tokens`: Tokens that were a cache-hit for that call.
- `completion_tokens_details`: Object containing reasoning_tokens.
- **ANTHROPIC_ONLY**: `cache_creation_input_tokens` are the number of tokens that were written to cache. (Anthropic charges for this).

## Quick Start

Note: OpenAI caching is only available for prompts containing 1024 tokens or more

**SDK**

```python
from litellm import completion 
import os

os.environ["OPENAI_API_KEY"] = ""

for _ in range(2):
    response = completion(
        model="gpt-5.6-terra",
        messages=[
            # System Message
            {
                "role": "system",
                "content": [
                    {
                        "type": "text",
                        "text": "Here is the full text of a complex legal agreement"
                        * 400,
                    }
                ],
            },
            {
                "role": "user",
                "content": [
                    {
                        "type": "text",
                        "text": "What are the key terms and conditions in this agreement?",
                    }
                ],
            },
            {
                "role": "assistant",
                "content": "Certainly! the key terms and conditions are the following: the contract is 1 year long for $10/mo",
            },
            {
                "role": "user",
                "content": [
                    {
                        "type": "text",
                        "text": "What are the key terms and conditions in this agreement?",
                    }
                ],
            },
        ],
        temperature=0.2,
        max_tokens=10,
    )

print("response=", response)
print("response.usage=", response.usage)

assert "prompt_tokens_details" in response.usage
assert response.usage.prompt_tokens_details.cached_tokens > 0
```

**PROXY**

1. Setup config.yaml

```yaml
model_list:
    - model_name: gpt-5.6-terra
      litellm_params:
        model: openai/gpt-5.6-terra
        api_key: os.environ/OPENAI_API_KEY
```

2. Start proxy

```bash
litellm --config /path/to/config.yaml
```

3. Test it!

```python
from openai import OpenAI
import os

client = OpenAI(
    api_key="LITELLM_PROXY_KEY", # sk-<your-litellm-api-key>
    base_url="LITELLM_PROXY_BASE" # http://0.0.0.0:4000
)

for _ in range(2):
    response = client.chat.completions.create(
        model="gpt-5.6-terra",
        messages=[
            # System Message
            {
                "role": "system",
                "content": [
                    {
                        "type": "text",
                        "text": "Here is the full text of a complex legal agreement"
                        * 400,
                    }
                ],
            },
            {
                "role": "user",
                "content": [
                    {
                        "type": "text",
                        "text": "What are the key terms and conditions in this agreement?",
                    }
                ],
            },
            {
                "role": "assistant",
                "content": "Certainly! the key terms and conditions are the following: the contract is 1 year long for $10/mo",
            },
            {
                "role": "user",
                "content": [
                    {
                        "type": "text",
                        "text": "What are the key terms and conditions in this agreement?",
                    }
                ],
            },
        ],
        temperature=0.2,
        max_tokens=10,
    )

print("response=", response)
print("response.usage=", response.usage)

assert "prompt_tokens_details" in response.usage
assert response.usage.prompt_tokens_details.cached_tokens > 0
```

### OpenAI `prompt_cache_key` and `prompt_cache_retention`

OpenAI prompt caching is [**automatic**](https://platform.openai.com/docs/guides/prompt-caching); no `cache_control` message annotations are needed. Any request with 1024+ prompt tokens is eligible for caching.

OpenAI also supports two optional parameters for more control over caching behavior:

- **`prompt_cache_key`** (string): a routing hint that improves cache hit rates for requests sharing long common prefixes. Requests with the same cache key are routed to the same backend, increasing the likelihood of a cache hit.
- **`prompt_cache_retention`** (`"in_memory"` or `"24h"`): controls cache TTL. Default is `"in_memory"` (5–10 min). Set to `"24h"` for extended caching that offloads KV tensors to GPU-local storage.

**SDK**

```python
from litellm import completion
import os

os.environ["OPENAI_API_KEY"] = ""

response = completion(
    model="gpt-5.6-terra",
    messages=[
        {
            "role": "system",
            "content": "You are an AI assistant tasked with analyzing legal documents. "
            + "Here is the full text of a complex legal agreement " * 400,
        },
        {
            "role": "user",
            "content": "What are the key terms and conditions?",
        },
    ],
    prompt_cache_key="legal-doc-analysis",
    prompt_cache_retention="24h",
)
print(response.usage)
```

**PROXY**

```python
from openai import OpenAI

client = OpenAI(
    api_key="LITELLM_PROXY_KEY",
    base_url="LITELLM_PROXY_BASE",
)

response = client.chat.completions.create(
    model="gpt-5.6-terra",
    messages=[
        {
            "role": "system",
            "content": "You are an AI assistant tasked with analyzing legal documents. "
            + "Here is the full text of a complex legal agreement " * 400,
        },
        {
            "role": "user",
            "content": "What are the key terms and conditions?",
        },
    ],
    extra_body={
        "prompt_cache_key": "legal-doc-analysis",
        "prompt_cache_retention": "24h",
    },
)
print(response.usage)
```

### OpenAI explicit breakpoints (GPT-5.6 and newer)

GPT-5.6 and newer also accept [explicit cache breakpoints](https://developers.openai.com/api/docs/guides/prompt-caching#prompt-cache-breakpoints): a `prompt_cache_breakpoint` marker on a content block plus a request-level `prompt_cache_options` that picks the mode (`implicit` keeps OpenAI's automatic breakpoint on the latest message alongside yours, `explicit` uses only yours) and the cache `ttl` (`30m`). LiteLLM passes both through on `/chat/completions`, `/responses` and, for Anthropic-shaped clients, `/v1/messages`. Models that accept the marker carry `supports_prompt_cache_breakpoint: true` in the cost map, and a GPT-5.6 or newer OpenAI model name the map has not flagged yet is treated the same way. To have LiteLLM place the marker for you from the deployment config, see the [auto-inject tutorial](../tutorials/prompt_caching.md#openai-gpt-56-and-newer)

**SDK**

```python
from litellm import completion
import os

os.environ["OPENAI_API_KEY"] = ""

response = completion(
    model="openai/gpt-5.6",
    messages=[
        {
            "role": "system",
            "content": [
                {
                    "type": "text",
                    "text": "You are an AI assistant tasked with analyzing legal documents. "
                    + "Here is the full text of a complex legal agreement " * 400,
                    "prompt_cache_breakpoint": {"mode": "explicit"},
                }
            ],
        },
        {
            "role": "user",
            "content": "What are the key terms and conditions?",
        },
    ],
    prompt_cache_options={"mode": "explicit", "ttl": "30m"},
)
print(response.usage.prompt_tokens_details)
```

**PROXY**

```python
from openai import OpenAI

client = OpenAI(
    api_key="LITELLM_PROXY_KEY",
    base_url="LITELLM_PROXY_BASE",
)

response = client.responses.create(
    model="gpt-5.6",
    input=[
        {
            "type": "message",
            "role": "developer",
            "content": [
                {
                    "type": "input_text",
                    "text": "You are an AI assistant tasked with analyzing legal documents. "
                    + "Here is the full text of a complex legal agreement " * 400,
                    "prompt_cache_breakpoint": {"mode": "explicit"},
                }
            ],
        },
        {
            "type": "message",
            "role": "user",
            "content": [{"type": "input_text", "text": "What are the key terms and conditions?"}],
        },
    ],
    extra_body={"prompt_cache_options": {"mode": "explicit", "ttl": "30m"}},
)
print(response.usage.input_tokens_details)
```

### Anthropic Example 

Anthropic charges for cache writes. 

Specify the content to cache with `"cache_control": {"type": "ephemeral"}`.

This same format also works for [Gemini / Vertex AI](#google-ai-studio--vertex-ai-gemini-example). For other providers, it will be ignored.

**SDK**

```python 
from litellm import completion 
import litellm 
import os 

litellm.set_verbose = True # 👈 SEE RAW REQUEST
os.environ["ANTHROPIC_API_KEY"] = "" 

response = completion(
    model="anthropic/claude-sonnet-5",
    messages=[
        {
            "role": "system",
            "content": [
                {
                    "type": "text",
                    "text": "You are an AI assistant tasked with analyzing legal documents.",
                },
                {
                    "type": "text",
                    "text": "Here is the full text of a complex legal agreement" * 400,
                    "cache_control": {"type": "ephemeral"},
                },
            ],
        },
        {
            "role": "user",
            "content": "what are the key terms and conditions in this agreement?",
        },
    ]
)

print(response.usage)
```
**PROXY**

1. Setup config.yaml

```yaml
model_list:
    - model_name: claude-sonnet-5
      litellm_params:
        model: anthropic/claude-sonnet-5
        api_key: os.environ/ANTHROPIC_API_KEY
```

2. Start proxy 

```bash
litellm --config /path/to/config.yaml
```

3. Test it! 

```python 
from openai import OpenAI 
import os

client = OpenAI(
    api_key="LITELLM_PROXY_KEY", # sk-<your-litellm-api-key>
    base_url="LITELLM_PROXY_BASE" # http://0.0.0.0:4000
)

response = client.chat.completions.create(
    model="claude-sonnet-5",
    messages=[
        {
            "role": "system",
            "content": [
                {
                    "type": "text",
                    "text": "You are an AI assistant tasked with analyzing legal documents.",
                },
                {
                    "type": "text",
                    "text": "Here is the full text of a complex legal agreement" * 400,
                    "cache_control": {"type": "ephemeral"},
                },
            ],
        },
        {
            "role": "user",
            "content": "what are the key terms and conditions in this agreement?",
        },
    ]
)

print(response.usage)
```

:::tip[Minimum tokens (Anthropic)]
Prompts below the minimum are processed without caching, and no error is returned. Check `cache_creation_input_tokens` in the response.

| Model | Min tokens |
|---|---|
| Claude Opus 5, Fable 5, Mythos 5 | 512 |
| Claude Sonnet 5, Opus 4.8 | 1,024 |
| Claude Opus 4.7 | 2,048 |
| Claude Haiku 4.5, Opus 4.5, Opus 4.6 | 4,096 |
| Claude Sonnet 4, Sonnet 4.5, Sonnet 4.6, Opus 4, Opus 4.1 | 1,024 |
| Claude 3.5 Haiku | 2,048 |
| Claude 3.x Haiku, Sonnet, Opus, 3.5 Sonnet, 3.7 Sonnet | 1,024 |

See [Anthropic's prompt caching docs](https://docs.anthropic.com/en/docs/build-with-claude/prompt-caching) for the full list; these minimums apply on every platform where each model is available.
:::

### Bedrock Example

LiteLLM automatically translates OpenAI-format `cache_control` markers to Bedrock's native `cachePoint` format, so no changes are needed to your existing code if you're already using `cache_control`.

:::tip[Minimum tokens (Bedrock)]
Prompts below the minimum are processed without caching, and no error is returned. Check `cache_creation_input_tokens` in the response.

| Model family | Min tokens per request |
|---|---|
| Claude Opus 5 | 512 |
| Claude Sonnet 5, Opus 4.8 | 1,024 |
| Claude Haiku 4.5, Opus 4.5, Opus 4.6, Opus 4.7 | 4,096 |
| Claude Sonnet 4.5, Sonnet 4.6 | 1,024 |
| Claude 3.5 Sonnet v2, Claude 3.7 Sonnet | 1,024 |

See [the Bedrock prompt caching docs](https://docs.aws.amazon.com/bedrock/latest/userguide/prompt-caching.html) for the full per-model table.
:::

**SDK**

```python
import litellm

response = litellm.completion(
    model="bedrock/us.anthropic.claude-sonnet-5",
    messages=[
        {
            "role": "system",
            "content": [
                {
                    "type": "text",
                    "text": "<your large system prompt here, at least 1,024 tokens on Claude Sonnet 4.x, 4,096 on Haiku 4.5 and Opus 4.5+>",
                    "cache_control": {"type": "ephemeral"}
                }
            ]
        },
        {"role": "user", "content": "What is prompt caching?"}
    ]
)

print(response.usage)
# cache_creation_input_tokens > 0 on first call (cache written)
# cache_read_input_tokens > 0 on subsequent calls (cache hit)
```

**PROXY**

1. Setup config.yaml

```yaml
model_list:
  - model_name: bedrock-claude-sonnet
    litellm_params:
      model: bedrock/us.anthropic.claude-sonnet-5
```

2. Start proxy

```bash
litellm --config /path/to/config.yaml
```

3. Test it!

```bash
curl -X POST http://localhost:4000/chat/completions \
  -H "Content-Type: application/json" \
  -H "Authorization: Bearer $LITELLM_MASTER_KEY" \
  -d '{
    "model": "bedrock-claude-sonnet",
    "messages": [
      {
        "role": "system",
        "content": [
          {
            "type": "text",
            "text": "<your large system prompt here, at least 1,024 tokens on Claude Sonnet 4.x, 4,096 on Haiku 4.5 and Opus 4.5+>",
            "cache_control": {"type": "ephemeral"}
          }
        ]
      },
      {"role": "user", "content": "What is prompt caching?"}
    ]
  }'
```

**Supported Bedrock models:**

| Model | Bedrock Model ID | Min Tokens | TTL Options |
|---|---|---|---|
| Claude Opus 5 | `anthropic.claude-opus-5` | 512 | 5 min, 1 hour |
| Claude Sonnet 5 | `anthropic.claude-sonnet-5` | 1,024 | 5 min, 1 hour |
| Claude Opus 4.8 | `anthropic.claude-opus-4-8` | 1,024 | 5 min, 1 hour |
| Claude 3.5 Sonnet v2 | `anthropic.claude-3-5-sonnet-20241022-v2:0` | 1,024 | 5 min, 1 hour |
| Claude 3.7 Sonnet | `anthropic.claude-3-7-sonnet-20250219-v1:0` | 1,024 | 5 min, 1 hour |
| Claude Opus 4 | `anthropic.claude-opus-4-20250514-v1:0` | 1,024 | 5 min, 1 hour |
| Claude Sonnet 4.5, 4.6 | `us.anthropic.claude-sonnet-4-5-*`, `us.anthropic.claude-sonnet-4-6-*` | 1,024 | 5 min, 1 hour |
| Claude Haiku 4.5 | `us.anthropic.claude-haiku-4-5-*` | 4,096 | 5 min, 1 hour |
| Claude Opus 4.5, 4.6, 4.7 | `us.anthropic.claude-opus-4-5-*`, `us.anthropic.claude-opus-4-6-*`, `us.anthropic.claude-opus-4-7-*` | 4,096 | 5 min, 1 hour |

Cross-region inference profiles are also supported for the models above.

See the [AWS Bedrock prompt caching docs](https://docs.aws.amazon.com/bedrock/latest/userguide/prompt-caching.html) for the full list of supported models and regions.

### Google AI Studio / Vertex AI (Gemini) Example

Use the same Anthropic-style `cache_control` format; LiteLLM automatically translates it to Google's [context caching API](https://ai.google.dev/api/caching).

**How it works under the hood:**
1. Messages with `cache_control` are separated and sent to Google's `cachedContents` API
2. The cached content ID is then passed as `cachedContent` in the Gemini request body
3. Works across all three providers: `gemini/` (Google AI Studio), `vertex_ai/`, and `vertex_ai_beta/`
4. Requires a minimum of **1024 tokens** in the cached content. Below that, caching is silently skipped

:::warning

Gemini 2.5 and newer already cache repeated prefixes implicitly, for free. Sending `cache_control` (by hand or through `cache_control_injection_points`) opts the request into explicit caching instead: the marked prefix is moved into a `cachedContents` object, so implicit caching no longer applies to it, and Google charges storage for each cache object. Only add `cache_control` for Gemini when you want explicit caching on purpose. See the [auto-inject tutorial](../tutorials/prompt_caching.md)

:::

**SDK**

```python
from litellm import completion
import os

os.environ["GEMINI_API_KEY"] = ""

response = completion(
    model="gemini/gemini-3.8-flash",
    messages=[
        {
            "role": "system",
            "content": [
                {
                    "type": "text",
                    "text": "You are an AI assistant tasked with analyzing legal documents.",
                },
                {
                    "type": "text",
                    "text": "Here is the full text of a complex legal agreement" * 400,
                    "cache_control": {"type": "ephemeral"},
                },
            ],
        },
        {
            "role": "user",
            "content": "what are the key terms and conditions in this agreement?",
        },
    ],
)

print(response.usage)
```
**PROXY**

1. Setup config.yaml

```yaml
model_list:
    - model_name: gemini-3.8-flash
      litellm_params:
        model: gemini/gemini-3.8-flash
        api_key: os.environ/GEMINI_API_KEY
```

2. Start proxy

```bash
litellm --config /path/to/config.yaml
```

3. Test it!

```python
from openai import OpenAI

client = OpenAI(
    api_key="LITELLM_PROXY_KEY",  # sk-<your-litellm-api-key>
    base_url="LITELLM_PROXY_BASE",  # http://0.0.0.0:4000
)

response = client.chat.completions.create(
    model="gemini-3.8-flash",
    messages=[
        {
            "role": "system",
            "content": [
                {
                    "type": "text",
                    "text": "You are an AI assistant tasked with analyzing legal documents.",
                },
                {
                    "type": "text",
                    "text": "Here is the full text of a complex legal agreement" * 400,
                    "cache_control": {"type": "ephemeral"},
                },
            ],
        },
        {
            "role": "user",
            "content": "what are the key terms and conditions in this agreement?",
        },
    ],
)

print(response.usage)
```

#### Vertex AI

For Vertex AI, use `vertex_ai/` prefix:

**SDK**

```python
from litellm import completion

response = completion(
    model="vertex_ai/gemini-3.8-flash",
    vertex_project="my-gcp-project",
    vertex_location="us-central1",
    messages=[
        {
            "role": "system",
            "content": [
                {
                    "type": "text",
                    "text": "You are an AI assistant tasked with analyzing legal documents.",
                },
                {
                    "type": "text",
                    "text": "Here is the full text of a complex legal agreement" * 400,
                    "cache_control": {"type": "ephemeral"},
                },
            ],
        },
        {
            "role": "user",
            "content": "what are the key terms and conditions in this agreement?",
        },
    ],
)

print(response.usage)
```
**PROXY**

1. Setup config.yaml

```yaml
model_list:
    - model_name: gemini-3.8-flash
      litellm_params:
        model: vertex_ai/gemini-3.8-flash
        vertex_project: my-gcp-project
        vertex_location: us-central1
```

2. Start proxy

```bash
litellm --config /path/to/config.yaml
```

3. Test it!

```python
from openai import OpenAI

client = OpenAI(
    api_key="LITELLM_PROXY_KEY",  # sk-<your-litellm-api-key>
    base_url="LITELLM_PROXY_BASE",  # http://0.0.0.0:4000
)

response = client.chat.completions.create(
    model="gemini-3.8-flash",
    messages=[
        {
            "role": "system",
            "content": [
                {
                    "type": "text",
                    "text": "You are an AI assistant tasked with analyzing legal documents.",
                },
                {
                    "type": "text",
                    "text": "Here is the full text of a complex legal agreement" * 400,
                    "cache_control": {"type": "ephemeral"},
                },
            ],
        },
        {
            "role": "user",
            "content": "what are the key terms and conditions in this agreement?",
        },
    ],
)

print(response.usage)
```

### Deepeek Example 

Works the same as OpenAI. 

```python 
from litellm import completion 
import litellm
import os 

os.environ["DEEPSEEK_API_KEY"] = "" 

litellm.set_verbose = True # 👈 SEE RAW REQUEST

model_name = "deepseek/deepseek-chat"
messages_1 = [
    {
        "role": "system",
        "content": "You are a history expert. The user will provide a series of questions, and your answers should be concise and start with `Answer:`",
    },
    {
        "role": "user",
        "content": "In what year did Qin Shi Huang unify the six states?",
    },
    {"role": "assistant", "content": "Answer: 221 BC"},
    {"role": "user", "content": "Who was the founder of the Han Dynasty?"},
    {"role": "assistant", "content": "Answer: Liu Bang"},
    {"role": "user", "content": "Who was the last emperor of the Tang Dynasty?"},
    {"role": "assistant", "content": "Answer: Li Zhu"},
    {
        "role": "user",
        "content": "Who was the founding emperor of the Ming Dynasty?",
    },
    {"role": "assistant", "content": "Answer: Zhu Yuanzhang"},
    {
        "role": "user",
        "content": "Who was the founding emperor of the Qing Dynasty?",
    },
]

message_2 = [
    {
        "role": "system",
        "content": "You are a history expert. The user will provide a series of questions, and your answers should be concise and start with `Answer:`",
    },
    {
        "role": "user",
        "content": "In what year did Qin Shi Huang unify the six states?",
    },
    {"role": "assistant", "content": "Answer: 221 BC"},
    {"role": "user", "content": "Who was the founder of the Han Dynasty?"},
    {"role": "assistant", "content": "Answer: Liu Bang"},
    {"role": "user", "content": "Who was the last emperor of the Tang Dynasty?"},
    {"role": "assistant", "content": "Answer: Li Zhu"},
    {
        "role": "user",
        "content": "Who was the founding emperor of the Ming Dynasty?",
    },
    {"role": "assistant", "content": "Answer: Zhu Yuanzhang"},
    {"role": "user", "content": "When did the Shang Dynasty fall?"},
]

response_1 = litellm.completion(model=model_name, messages=messages_1)
response_2 = litellm.completion(model=model_name, messages=message_2)

# Add any assertions here to check the response
print(response_2.usage)
```

## Calculate Cost 

Cost cache-hit prompt tokens can differ from cache-miss prompt tokens.

Use the `completion_cost()` function for calculating cost ([handles prompt caching cost calculation](https://github.com/BerriAI/litellm/blob/f7ce1173f3315cc6cae06cf9bcf12e54a2a19705/litellm/llms/anthropic/cost_calculation.py#L12) as well). [**See more helper functions**](./token_usage.md)

```python
cost = completion_cost(completion_response=response, model=model)
```

### Usage

**SDK**

```python
from litellm import completion, completion_cost
import litellm 
import os 

litellm.set_verbose = True # 👈 SEE RAW REQUEST
os.environ["ANTHROPIC_API_KEY"] = "" 
model = "anthropic/claude-sonnet-5"
response = completion(
    model=model,
    messages=[
        {
            "role": "system",
            "content": [
                {
                    "type": "text",
                    "text": "You are an AI assistant tasked with analyzing legal documents.",
                },
                {
                    "type": "text",
                    "text": "Here is the full text of a complex legal agreement" * 400,
                    "cache_control": {"type": "ephemeral"},
                },
            ],
        },
        {
            "role": "user",
            "content": "what are the key terms and conditions in this agreement?",
        },
    ]
)

print(response.usage)

cost = completion_cost(completion_response=response, model=model) 

formatted_string = f"${float(cost):.10f}"
print(formatted_string)
```
**PROXY**

LiteLLM returns the calculated cost in the response headers - `x-litellm-response-cost` 

```python
from openai import OpenAI

client = OpenAI(
    api_key="LITELLM_PROXY_KEY", # sk-<your-litellm-api-key>..
    base_url="LITELLM_PROXY_BASE" # http://0.0.0.0:4000
)
response = client.chat.completions.with_raw_response.create(
    messages=[{
        "role": "user",
        "content": "Say this is a test",
    }],
    model="gpt-5.6-luna",
)
print(response.headers.get('x-litellm-response-cost'))

completion = response.parse()  # get the object that `chat.completions.create()` would have returned
print(completion)
```

## Check Model Support

Check if a model supports prompt caching with `supports_prompt_caching()` 

**SDK**

```python
from litellm.utils import supports_prompt_caching

supports_pc: bool = supports_prompt_caching(model="anthropic/claude-sonnet-5")

assert supports_pc
```

**PROXY**

Use the `/model/info` endpoint to check if a model on the proxy supports prompt caching 

1. Setup config.yaml 

```yaml
model_list:
    - model_name: claude-sonnet-5
      litellm_params:
        model: anthropic/claude-sonnet-5
        api_key: os.environ/ANTHROPIC_API_KEY
```

2. Start proxy 

```bash
litellm --config /path/to/config.yaml
```

3. Test it! 

```bash
curl -L -X GET 'http://0.0.0.0:4000/v1/model/info' \
-H "Authorization: Bearer $LITELLM_API_KEY" \
```

**Expected Response**

```bash
{
    "data": [
        {
            "model_name": "claude-sonnet-5",
            "litellm_params": {
                "model": "anthropic/claude-sonnet-5"
            },
            "model_info": {
                "key": "claude-sonnet-5",
                ...
                "supports_prompt_caching": true # 👈 LOOK FOR THIS!
            }
        }
    ]
}
```

This checks our maintained [model info/cost map](https://github.com/BerriAI/litellm/blob/main/model_prices_and_context_window.json)

## Read More

:::tip[Auto-Inject Prompt Caching]
Want LiteLLM to automatically add `cache_control` directives without modifying your code? 

See [**Auto-Inject Prompt Caching Tutorial**](../tutorials/prompt_caching.md) to learn how to use `cache_control_injection_points` to automatically cache system messages, specific messages by index, or custom injection patterns.
:::

## Related pages

- [Trimming Input Messages](https://docs.litellm.ai/docs/completion/message_trimming.md)
- [Prompt Formatting](https://docs.litellm.ai/docs/completion/prompt_formatting.md)
