---
title: "Input Params"
url: "/docs/completion/input"
canonical_url: "https://docs.litellm.ai/docs/completion/input"
type: "docs"
last_updated: "2026-10-08"
related:
  - "/docs/learn/sdk_quickstart"
  - "/docs/embedding/supported_embedding"
---
# Input Params

> Index of all LiteLLM docs: https://docs.litellm.ai/llms.txt


## Common Params 
LiteLLM accepts and translates the [OpenAI Chat Completion params](https://platform.openai.com/docs/api-reference/chat/create) across all providers. 

### Usage
```python
import litellm

# set env variables
os.environ["OPENAI_API_KEY"] = "your-openai-key"

## SET MAX TOKENS - via completion() 
response = litellm.completion(
            model="gpt-5.6-luna",
            messages=[{ "content": "Hello, how are you?","role": "user"}],
            max_tokens=10
        )

print(response)
```

### Translated OpenAI params

Use this function to get an up-to-date list of supported openai params for any model + provider. 

```python
from litellm import get_supported_openai_params

response = get_supported_openai_params(model="anthropic.claude-sonnet-5", custom_llm_provider="bedrock")

print(response) # ["max_tokens", "tools", "tool_choice", "stream"]
```

This table is the output of `litellm.get_supported_openai_params()` for the model shown in each row. Support is model dependent within a provider (for example Bedrock Llama models do not list `tools` or `tool_choice`), so call the function for the exact model you use

`stream_options`, `extra_headers` and `max_retries` are not checked against this list and are accepted for every provider, and `stream_options={"include_usage": True}` returns usage on the final chunk for every provider

| Provider | Model checked | temperature | max_completion_tokens | max_tokens | top_p | stream | stop | n | presence_penalty | frequency_penalty | functions | function_call | logit_bias | user | response_format | seed | tools | tool_choice | logprobs | top_logprobs |
|---|---|---|---|---|---|---|---|---|---|---|---|---|---|---|---|---|---|---|---|---|
| Anthropic | `claude-sonnet-4-5-20250929` | ✅ | ✅ | ✅ | ✅ | ✅ | ✅ |  |  |  |  |  |  | ✅ | ✅ |  | ✅ | ✅ |  |  |
| OpenAI | `gpt-4o` | ✅ | ✅ | ✅ | ✅ | ✅ | ✅ | ✅ | ✅ | ✅ | ✅ | ✅ | ✅ | ✅ | ✅ | ✅ | ✅ | ✅ | ✅ | ✅ |
| Azure OpenAI | `gpt-4o` | ✅ | ✅ | ✅ | ✅ | ✅ | ✅ | ✅ | ✅ | ✅ | ✅ | ✅ | ✅ | ✅ | ✅ | ✅ | ✅ | ✅ | ✅ | ✅ |
| xAI | `grok-3` | ✅ | ✅ | ✅ | ✅ | ✅ | ✅ | ✅ | ✅ | ✅ |  |  | ✅ | ✅ | ✅ | ✅ | ✅ | ✅ | ✅ | ✅ |
| Replicate | `meta/llama-2-70b-chat` | ✅ |  | ✅ | ✅ | ✅ | ✅ |  |  |  | ✅ | ✅ |  |  |  | ✅ | ✅ | ✅ |  |  |
| Anyscale | `meta-llama/Llama-2-70b-chat-hf` | ✅ |  | ✅ | ✅ | ✅ | ✅ |  | ✅ | ✅ |  |  |  |  |  |  |  |  |  |  |
| Cohere | `command-r` | ✅ | ✅ | ✅ | ✅ | ✅ | ✅ | ✅ | ✅ | ✅ |  |  |  |  |  | ✅ | ✅ | ✅ |  |  |
| Huggingface | `meta-llama/Llama-3.1-8B-Instruct` | ✅ | ✅ | ✅ | ✅ | ✅ | ✅ | ✅ | ✅ | ✅ | ✅ | ✅ | ✅ |  | ✅ | ✅ | ✅ | ✅ | ✅ | ✅ |
| Openrouter | `openai/gpt-4o` | ✅ | ✅ | ✅ | ✅ | ✅ | ✅ | ✅ | ✅ | ✅ | ✅ | ✅ | ✅ |  | ✅ | ✅ | ✅ | ✅ | ✅ | ✅ |
| AI21 | `jamba-1.5-large` | ✅ | ✅ | ✅ |  | ✅ | ✅ | ✅ |  |  |  |  |  |  | ✅ | ✅ | ✅ | ✅ |  |  |
| VertexAI | `gemini-2.5-flash` | ✅ | ✅ | ✅ | ✅ | ✅ | ✅ | ✅ | ✅ | ✅ | ✅ |  |  |  | ✅ | ✅ | ✅ | ✅ | ✅ | ✅ |
| Bedrock | `anthropic.claude-3-5-sonnet-20240620-v1:0` | ✅ | ✅ | ✅ | ✅ | ✅ | ✅ |  |  |  |  |  |  |  | ✅ |  | ✅ | ✅ |  |  |
| Sagemaker | `jumpstart-dft-meta-textgeneration-llama-2-7b` | ✅ | ✅ | ✅ | ✅ | ✅ | ✅ | ✅ |  |  |  |  |  |  |  |  |  |  |  |  |
| TogetherAI | `meta-llama/Llama-3-70b-chat-hf` | ✅ | ✅ | ✅ | ✅ | ✅ | ✅ | ✅ | ✅ | ✅ | ✅ | ✅ | ✅ |  | ✅ | ✅ | ✅ | ✅ | ✅ | ✅ |
| Sambanova | `Meta-Llama-3.1-8B-Instruct` | ✅ | ✅ | ✅ | ✅ | ✅ | ✅ |  |  |  |  |  |  |  | ✅ |  |  |  |  |  |
| AlephAlpha | `luminous-base` | ✅ |  | ✅ | ✅ | ✅ | ✅ | ✅ | ✅ | ✅ |  |  |  |  |  |  |  |  |  |  |
| NLP Cloud | `dolphin` | ✅ |  | ✅ | ✅ | ✅ | ✅ | ✅ | ✅ | ✅ |  |  |  |  |  |  |  |  |  |  |
| Petals | `petals-team/StableBeluga2` | ✅ |  | ✅ | ✅ | ✅ |  |  |  |  |  |  |  |  |  |  |  |  |  |  |
| Ollama | `llama3` | ✅ | ✅ | ✅ | ✅ | ✅ | ✅ |  |  | ✅ |  |  |  |  | ✅ | ✅ |  |  |  |  |
| Databricks | `databricks-dbrx-instruct` | ✅ | ✅ | ✅ | ✅ | ✅ | ✅ | ✅ |  |  |  |  |  |  | ✅ |  | ✅ | ✅ |  |  |
| ClarifAI | `openai.chat-completion.gpt-4o` | ✅ | ✅ | ✅ | ✅ | ✅ |  |  | ✅ | ✅ |  |  |  |  | ✅ |  | ✅ | ✅ |  |  |
| Github | `gpt-4o` | ✅ | ✅ | ✅ | ✅ | ✅ | ✅ | ✅ | ✅ | ✅ | ✅ | ✅ | ✅ | ✅ | ✅ | ✅ | ✅ | ✅ | ✅ | ✅ |
| Novita AI | `meta-llama/llama-3-8b-instruct` | ✅ | ✅ | ✅ | ✅ | ✅ | ✅ | ✅ | ✅ | ✅ | ✅ | ✅ | ✅ |  | ✅ | ✅ | ✅ | ✅ | ✅ | ✅ |
| Bytez | `google/gemma-3-1b-it` | ✅ | ✅ | ✅ | ✅ | ✅ |  | ✅ |  |  |  |  |  |  |  |  |  |  |  |  |
| OVHCloud AI Endpoints | `Meta-Llama-3_3-70B-Instruct` | ✅ | ✅ | ✅ | ✅ | ✅ | ✅ | ✅ | ✅ | ✅ | ✅ | ✅ | ✅ |  | ✅ | ✅ | ✅ | ✅ | ✅ | ✅ |

:::note

By default, LiteLLM raises an exception if the openai param being passed in isn't supported. 

To drop the param instead, set `litellm.drop_params = True` or `completion(..drop_params=True)`.

This **ONLY DROPS UNSUPPORTED OPENAI PARAMS**. 

LiteLLM assumes any non-openai param is provider specific and passes it in as a kwarg in the request body

::: 

## Input Params

```python
def completion(
    model: str,
    messages: list = [],
    timeout: float | str | httpx.Timeout | None = None,
    temperature: float | None = None,
    top_p: float | None = None,
    n: int | None = None,
    stream: bool | None = None,
    stream_options: dict | None = None,
    stop=None,
    max_completion_tokens: int | None = None,
    max_tokens: int | None = None,
    modalities: list[ChatCompletionModality] | None = None,
    prediction: ChatCompletionPredictionContentParam | None = None,
    audio: ChatCompletionAudioParam | None = None,
    presence_penalty: float | None = None,
    frequency_penalty: float | None = None,
    logit_bias: dict | None = None,
    user: str | None = None,
    # openai v1.0+ new params
    reasoning_effort: Literal["none", "minimal", "low", "medium", "high", "xhigh", "max", "default"] | None = None,
    verbosity: Literal["low", "medium", "high"] | None = None,
    response_format: dict | type[BaseModel] | None = None,
    seed: int | None = None,
    tools: list | None = None,
    tool_choice: str | dict | None = None,
    logprobs: bool | None = None,
    top_logprobs: int | None = None,
    parallel_tool_calls: bool | None = None,
    web_search_options: OpenAIWebSearchOptions | None = None,
    include_server_side_tool_invocations: bool | None = None,
    deployment_id=None,
    extra_headers: dict | None = None,
    safety_identifier: str | None = None,
    service_tier: str | None = None,
    store: bool | None = None,
    prompt_cache_key: str | None = None,
    # soon to be deprecated params by OpenAI
    functions: list | None = None,
    function_call: str | None = None,
    # set api_base, api_version, api_key
    base_url: str | None = None,
    api_version: str | None = None,
    api_key: str | None = None,
    model_list: list | None = None,  # pass in a list of api_base,keys, etc.
    # Optional liteLLM function params
    thinking: AnthropicThinkingParam | None = None,
    # Session management
    shared_session: Optional["ClientSession"] = None,
    # Per-request JSON schema validation (overrides litellm.enable_json_schema_validation)
    enable_json_schema_validation: bool | None = None,
    **kwargs,
) -> ModelResponse | CustomStreamWrapper:
    ...
```
### Required Fields

- `model`: *string* - ID of the model to use. Refer to the model endpoint compatibility table for details on which models work with the Chat API.
  
- `messages`: *array* - A list of messages comprising the conversation so far.

#### Properties of `messages`
*Note* - Each message in the array contains the following properties:

- `role`: *string* - The role of the message's author. Roles can be: system, user, assistant, function or tool.

- `content`: *string or list[dict] or null* - The contents of the message. It is required for all messages, but may be null for assistant messages with function calls.

- `name`: *string (optional)* - The name of the author of the message. It is required if the role is "function". The name should match the name of the function represented in the content. It can contain characters (a-z, A-Z, 0-9), and underscores, with a maximum length of 64 characters.

- `function_call`: *object (optional)* - The name and arguments of a function that should be called, as generated by the model.

- `tool_call_id`: *str (optional)* - Tool call that this message is responding to.

[**See All Message Values**](https://github.com/BerriAI/litellm/blob/main/litellm/types/llms/openai.py#L664)

#### Content Types

`content` can be a string (text only) or a list of content blocks (multimodal):

| Type | Description | Docs |
|------|-------------|------|
| `text` | Text content | [Type Definition](https://github.com/BerriAI/litellm/blob/main/litellm/types/llms/openai.py#L598) |
| `image_url` | Images | [Vision](./vision.md) |
| `input_audio` | Audio input | [Audio](./audio.md) |
| `video_url` | Video input | [Type Definition](https://github.com/BerriAI/litellm/blob/main/litellm/types/llms/openai.py#L625) |
| `file` | Files | [Document Understanding](./document_understanding.md) |
| `document` | Documents/PDFs | [Document Understanding](./document_understanding.md) |

**Examples:**
```python
# Text
messages=[{"role": "user", "content": [{"type": "text", "text": "Hello!"}]}]

# Image
messages=[{"role": "user", "content": [{"type": "image_url", "image_url": {"url": "https://example.com/image.jpg"}}]}]

# Audio
messages=[{"role": "user", "content": [{"type": "input_audio", "input_audio": {"data": "<base64>", "format": "wav"}}]}]

# Video
messages=[{"role": "user", "content": [{"type": "video_url", "video_url": {"url": "https://example.com/video.mp4"}}]}]

# File
messages=[{"role": "user", "content": [{"type": "file", "file": {"file_id": "https://example.com/doc.pdf"}}]}]

# Document
messages=[{"role": "user", "content": [{"type": "document", "source": {"type": "text", "media_type": "application/pdf", "data": "<base64>"}}]}]

# Combining multiple types (multimodal)
messages=[{"role": "user", "content": [
    {"type": "text", "text": "Generate a product description based on this image"},
    {"type": "image_url", "image_url": {"url": "https://example.com/image.jpg"}}
]}]
```

## Optional Fields

- `temperature`: *number or null (optional)* - The sampling temperature to be used, between 0 and 2. Higher values like 0.8 produce more random outputs, while lower values like 0.2 make outputs more focused and deterministic. 

- `top_p`: *number or null (optional)* - An alternative to sampling with temperature. It instructs the model to consider the results of the tokens with top_p probability. For example, 0.1 means only the tokens comprising the top 10% probability mass are considered.

- `n`: *integer or null (optional)* - The number of chat completion choices to generate for each input message.

- `stream`: *boolean or null (optional)* - If set to true, it sends partial message deltas. Tokens will be sent as they become available, with the stream terminated by a [DONE] message.

- `stream_options` *dict or null (optional)* - Options for streaming response. Only set this when you set `stream: true`

    - `include_usage` *boolean (optional)* - If set, an additional chunk will be streamed before the data: [DONE] message. The usage field on this chunk shows the token usage statistics for the entire request, and the choices field will always be an empty array. All other chunks will also include a usage field, but with a null value. 

- `stop`: *string/ array/ null (optional)* - Up to 4 sequences where the API will stop generating further tokens.
  
  **Note**: OpenAI supports a maximum of 4 stop sequences. If you provide more than 4, LiteLLM will automatically truncate the list to the first 4 elements. To disable this automatic truncation, set `litellm.disable_stop_sequence_limit = True`.

- `max_completion_tokens`: *integer (optional)* -  An upper bound for the number of tokens that can be generated for a completion, including visible output tokens and reasoning tokens.

- `max_tokens`: *integer (optional)* - The maximum number of tokens to generate in the chat completion.

- `presence_penalty`: *number or null (optional)* - It is used to penalize new tokens based on their existence in the text so far.

- `response_format`: *dict or Pydantic model class (optional)* - Specifies the response format. Pass a Pydantic model class to request schema-based output; see [JSON mode and structured outputs](./json_mode.md).

    - Setting to `{ "type": "json_object" }` enables JSON mode, which guarantees the message the model generates is valid JSON.
    
    - Important: when using JSON mode, you must also instruct the model to produce JSON yourself via a system or user message. Without this, the model may generate an unending stream of whitespace until the generation reaches the token limit, resulting in a long-running and seemingly "stuck" request. Also note that the message content may be partially cut off if finish_reason="length", which indicates the generation exceeded max_tokens or the conversation exceeded the max context length.

- `seed`: *integer or null (optional)* - This feature is in Beta. If specified, our system will make a best effort to sample deterministically, such that repeated requests with the same seed and parameters should return the same result. Determinism is not guaranteed, and you should refer to the `system_fingerprint` response parameter to monitor changes in the backend.

- `tools`: *array (optional)* - A list of tools the model may call. Use this to provide a list of functions the model may generate JSON inputs for.

    - `type`: *string* - The type of the tool. You can set this to `"function"` or `"mcp"` (matching the `/responses` schema) to call LiteLLM-registered MCP servers directly from `/chat/completions`.

    - `function`: *object* - Required for function tools.

- `tool_choice`: *string or dict (optional)* - Controls which tool, if any, the model calls. Use `none` or `auto`, or provide a dict naming a specific tool.

    - `none` is the default when no functions are present. `auto` is the default if functions are present.

- `parallel_tool_calls`: *boolean (optional)* - Whether to enable parallel function calling during tool use. OpenAI default is true.

- `frequency_penalty`: *number or null (optional)* - It is used to penalize new tokens based on their frequency in the text so far.

- `logit_bias`: *map (optional)* - Used to modify the probability of specific tokens appearing in the completion.

- `user`: *string (optional)* - A unique identifier representing your end-user. This can help OpenAI to monitor and detect abuse.

- `timeout`: *float, string, `httpx.Timeout` or null (optional)* - Request timeout in seconds. A string can contain the numeric timeout value; use `httpx.Timeout` to configure timeout phases separately.

- `logprobs`: * bool (optional)* - Whether to return log probabilities of the output tokens or not. If true returns the log probabilities of each output token returned in the content of message
        
- `top_logprobs`: *int (optional)* - An integer between 0 and 5 specifying the number of most likely tokens to return at each token position, each with an associated log probability. `logprobs` must be set to true if this parameter is used.

- `safety_identifier`: *string (optional)* - A unique identifier for tracking and managing safety-related requests. This parameter helps with safety monitoring and compliance tracking.

- `headers`: *dict (optional)* - A dictionary of headers to be sent with the request.

- `extra_headers`: *dict (optional)* - Alternative to `headers`, used to send extra headers in LLM API request. 

- `reasoning_effort`: *string or null (optional)* - Sets the reasoning effort for supported models. Values are `none`, `minimal`, `low`, `medium`, `high`, `xhigh`, `max` and `default`. See [reasoning content](../reasoning_content.md).

- `verbosity`: *string or null (optional)* - Sets the output verbosity to `low`, `medium` or `high` for supported OpenAI GPT-5 family models.

- `thinking`: *object or null (optional)* - Anthropic-style thinking configuration. See [reasoning content](../reasoning_content.md).

- `modalities`: *array or null (optional)* - Requested output modalities, such as text and audio. See [audio](./audio.md).

- `audio`: *object or null (optional)* - Audio output settings, such as voice and format. See [audio](./audio.md).

- `prediction`: *object or null (optional)* - Expected output content that can reduce latency when much of the response is known in advance. See [predicted outputs](./predict_outputs.md).

- `web_search_options`: *object or null (optional)* - Options for supported built-in web search models and endpoints, such as search context size. See [web search](./web_search.md).

- `include_server_side_tool_invocations`: *boolean or null (optional)* - For supported Gemini/Vertex requests, sets `toolConfig.includeServerSideToolInvocations`.

- `service_tier`: *string or null (optional)* - Requests a processing service tier where the model supports it.

- `store`: *boolean or null (optional)* - Controls whether the provider stores the response where supported.

- `prompt_cache_key`: *string or null (optional)* - OpenAI prompt-cache routing hint for requests with shared prefixes. See [prompt caching](./prompt_caching.md).

- `shared_session`: *aiohttp `ClientSession` or null (optional)* - Reuses a session across asynchronous API calls. See [shared sessions](./shared_session.md).

- `enable_json_schema_validation`: *boolean or null (optional)* - Enables or disables per-request validation of generated JSON against the requested schema, overriding the global setting. See [JSON mode and structured outputs](./json_mode.md).

- `api_key`: *string or null (optional)* - Provider API key to use for this request.

- `base_url`: *string or null (optional)* - Provider API base URL. Alias for `api_base`.

- `model_list`: *list or null (optional)* - Model deployment configurations used to select and call deployments matching `model`.

#### Deprecated Params
- `functions`: *array* - A list of functions that the model may use to generate JSON inputs. Each function should have the following properties:

    - `name`: *string* - The name of the function to be called. It should contain a-z, A-Z, 0-9, underscores and dashes, with a maximum length of 64 characters.
    
    - `description`: *string (optional)* - A description explaining what the function does. It helps the model to decide when and how to call the function.
    
    - `parameters`: *object* - The parameters that the function accepts, described as a JSON Schema object.
    
- `function_call`: *string or object (optional)* - Controls how the model responds to function calls.

#### litellm-specific params 

- `api_base`: *string (optional)* - The api endpoint you want to call the model with

- `api_version`: *string (optional)* - (Azure-specific) the api version for the call

- `num_retries`: *int (optional)* - The number of times to retry the API call if an APIError, TimeoutError or ServiceUnavailableError occurs 

- `context_window_fallback_dict`: *dict (optional)* - A mapping of model to use if call fails due to context window error

- `fallbacks`: *list (optional)* - A list of model names + params to be used, in case the initial call fails

- `metadata`: *dict (optional)* - Any additional data you want to be logged when the call is made (sent to logging integrations, eg. promptlayer and accessible via custom callback function)

- `custom_llm_provider`: *string (optional)* - Provider identifier to use when resolving the model and its API configuration

- `drop_params`: *boolean (optional)* - Drops unsupported OpenAI parameters instead of raising an error. See [drop unsupported params](./drop_params.md).

- `additional_drop_params`: *list of strings (optional)* - OpenAI parameters to drop from the request. See [drop unsupported params](./drop_params.md).

- `allowed_openai_params`: *list of strings (optional)* - OpenAI parameters allowed through for this request. See [drop unsupported params](./drop_params.md).

- `mock_response`: *string or response object (optional)* - Returns a mock completion response without calling the model. See [mock requests](./mock_requests.md).

- `max_retries`: *integer (optional)* - Maximum number of retries for the API call

- `ssl_verify`: *boolean or string (optional)* - Enables or disables SSL verification, or sets a custom CA bundle path. See [SSL security settings](../guides/security_settings.md).

- `merge_reasoning_content_in_choices`: *boolean (optional)* - For streaming responses, adds `reasoning_content` to `content` inside `<think>` tags for clients that expect reasoning in content

- `prompt_id`: *string (optional)* - Managed prompt ID passed to prompt-management hooks. See [prompt management](../prompt_management.md).

- `prompt_variables`: *dict (optional)* - Values used to fill managed prompt variables when prompt-management hooks are configured. See [prompt management](../prompt_management.md).

- `litellm_system_prompt`: *string (optional)* - Prepends this prompt to the first system message, or adds a system message at the start if none exists

- `base_model`: *string (optional)* - Underlying model name used for provider-specific model and parameter configuration. See [fine-tuned models](../guides/finetuned_models.md).

- `supports_system_message`: *boolean (optional)* - Set to false when the model does not accept system-role messages; LiteLLM then remaps those messages

- `ensure_alternating_roles`: *boolean (optional)* - Adds user or assistant continuation messages where needed to alternate conversation roles. See [message sanitization](./message_sanitization.md).

- `user_continue_message`: *message object (optional)* - Custom user message inserted when alternating roles require a user turn. See [message sanitization](./message_sanitization.md).

- `assistant_continue_message`: *message object (optional)* - Custom assistant message inserted when alternating roles require an assistant turn. See [message sanitization](./message_sanitization.md).

- `cooldown_time`: *float (optional)* - Cooldown duration in seconds before a failed routed deployment can be retried

**CUSTOM MODEL COST** 
- `input_cost_per_token`: *float (optional)* - The cost per input token for the completion call 

- `output_cost_per_token`: *float (optional)* - The cost per output token for the completion call 

- `cost_per_second`: *float (optional)* - Per-second rate used to calculate costs for models billed by duration

- `input_cost_per_second`: *float (optional)* - Per-second input rate used to calculate costs for models billed by duration

- `output_cost_per_second`: *float (optional)* - Per-second output rate used to calculate costs for models billed by duration

**CUSTOM PROMPT TEMPLATE** (See [prompt formatting for more info](./prompt_formatting.md#format-prompt-yourself))
- `initial_prompt_value`: *string (optional)* - Initial string applied at the start of the input messages

- `roles`: *dict (optional)* - Dictionary specifying how to format the prompt based on the role + message passed in via `messages`. 

- `final_prompt_value`: *string (optional)* - Final string applied at the end of the input messages

- `bos_token`: *string (optional)* - Initial string applied at the start of a sequence

- `eos_token`: *string (optional)* - Initial string applied at the end of a sequence

- `hf_model_name`: *string (optional)* - [Sagemaker Only] The corresponding huggingface name of the model, used to pull the right chat template for the model.

## Related pages

- [SDK Quickstart](https://docs.litellm.ai/docs/learn/sdk_quickstart.md)
- [embedding()](https://docs.litellm.ai/docs/embedding/supported_embedding.md)
