Skip to main content

Web Search Integration

Enable transparent server-side web search execution for any LLM provider. LiteLLM automatically intercepts web search tool calls and executes them using your configured search provider (Parallel, Perplexity, Tavily, and others).

Quick Start​

1. Configure Web Search Interception​

Add to your config.yaml:

model_list:
- model_name: gpt-5.6-terra
litellm_params:
model: openai/gpt-5.6-terra
api_key: os.environ/OPENAI_API_KEY

litellm_settings:
callbacks: ["websearch_interception"]
websearch_interception_params:
enabled_providers:
- openai
- minimax
- anthropic
search_tool_name: perplexity-search # Optional

search_tools:
- search_tool_name: perplexity-search
litellm_params:
search_provider: perplexity
api_key: os.environ/PERPLEXITY_API_KEY

2. Use with Any Provider​

import litellm

response = await litellm.acompletion(
model="gpt-5.6-terra",
messages=[
{"role": "user", "content": "What's the weather in San Francisco today?"}
],
tools=[
{
"type": "function",
"function": {
"name": "litellm_web_search",
"description": "Search the web for information",
"parameters": {
"type": "object",
"properties": {
"query": {"type": "string", "description": "Search query"}
},
"required": ["query"]
}
}
}
]
)

# Response includes search results automatically!
print(response.choices[0].message.content)

How It Works​

When a model makes a web search tool call, LiteLLM:

  1. Detects the litellm_web_search tool call in the response
  2. Executes the search using your configured search provider
  3. Makes a follow-up request with the search results
  4. Returns the final answer to the user

Result: One API call from user → Complete answer with search results

The Search Tool the Model Sees​

Whatever web search tool the request carries, LiteLLM replaces it with its own litellm_web_search definition before the model sees it. Anthropic's web_search_20250305, the Responses API's web_search_preview, Claude Code's web_search, and a litellm_web_search function tool you define yourself with only a query parameter all reach the model as the schema below, in the tool format of the API you called (an Anthropic input_schema, a Chat Completions function.parameters, or a flat Responses function tool).

litellm_web_search input schema
{
"type": "object",
"properties": {
"query": {
"type": "string",
"description": "The search query to execute"
},
"objective": {
"type": "string",
"description": "Natural-language description of the goal behind the search, including any source or freshness requirements."
},
"search_queries": {
"type": "array",
"items": {"type": "string"},
"description": "Two to five short keyword queries (3-6 words each) covering different angles of the objective, e.g. varying names, synonyms, or phrasings. Provide together with objective for the best results."
}
},
"required": ["query"]
}

query is required and is what most search providers receive. objective and search_queries are optional: they let the model say what it is after and fan out several keyword searches in one tool call. LiteLLM forwards them only to search providers whose API takes that shape natively, which today means Parallel AI. There the search_queries list becomes the provider's queries and objective goes alongside it, unless the search tool's litellm_params already sets an objective, which is kept over the model's. Every other search provider (Perplexity, Tavily, Exa, and the rest) keeps receiving the single query string, and a model that fills only query behaves exactly as before on every provider.

The optional fields are validated before they are forwarded: a blank objective is ignored, search_queries has to be an array (a bare string is ignored rather than split into characters), entries that are not non-empty strings are dropped, and only the first five queries are kept, matching Parallel's own cap.

With a Parallel AI search tool configured, a tool call from the model and the request LiteLLM builds from it look like this. The outbound body is what --detailed_debug prints as the request sent to https://api.parallel.ai/v1/search.

Tool call emitted by the model
{
"name": "litellm_web_search",
"input": {
"query": "latest stable Node.js release",
"objective": "Find the most current stable Node.js release version and what changes were included in that release",
"search_queries": ["latest stable Node.js release", "Node.js newest version changelog", "current Node.js LTS release"]
}
}
Request LiteLLM sends to Parallel AI
{
"objective": "Find the most current stable Node.js release version and what changes were included in that release",
"search_queries": ["latest stable Node.js release", "Node.js newest version changelog", "current Node.js LTS release"],
"mode": "basic",
"advanced_settings": {"max_results": 5}
}

The same tool call with a Tavily search tool sends Tavily "query": "latest stable Node.js release" and nothing from the other two fields. For clients that sent an Anthropic-native web_search_* tool, the server_tool_use block in the final response still shows only query.

Supported Providers​

Web search integration works with all providers that use:

  • ✅ Base HTTP Handler (BaseLLMHTTPHandler)
  • ✅ OpenAI Completion Handler (OpenAIChatCompletion)

Providers Using Base HTTP Handler​

ProviderStatusNotes
OpenAI✅ SupportedGPT-4, GPT-3.5, etc.
Anthropic✅ SupportedClaude models via HTTP handler
MiniMax✅ SupportedAll MiniMax models
Mistral✅ SupportedMistral AI models
Cohere✅ SupportedCommand models
Fireworks AI✅ SupportedAll Fireworks models
Together AI✅ SupportedAll Together AI models
Groq✅ SupportedAll Groq models
Perplexity✅ SupportedPerplexity models
DeepSeek✅ SupportedDeepSeek models
xAI✅ SupportedGrok models
Hugging Face✅ SupportedInference API models
OCI✅ SupportedOracle Cloud models
Vertex AI✅ SupportedGoogle Vertex AI models
Bedrock✅ SupportedAWS Bedrock models (converse_like route)
Azure OpenAI✅ SupportedAzure-hosted OpenAI models
Sagemaker✅ SupportedAWS Sagemaker models
Databricks✅ SupportedDatabricks models
DataRobot✅ SupportedDataRobot models
Hosted VLLM✅ SupportedSelf-hosted VLLM
Heroku✅ SupportedHeroku-hosted models
RAGFlow✅ SupportedRAGFlow models
Compactif✅ SupportedCompactif models
Cometapi✅ SupportedComet API models
A2A✅ SupportedAgent-to-Agent models
Bytez✅ SupportedBytez models

Providers Using OpenAI Handler​

ProviderStatusNotes
OpenAI✅ SupportedNative OpenAI API
Azure OpenAI✅ SupportedAzure-hosted OpenAI
OpenAI-Compatible✅ SupportedAny OpenAI-compatible API

Configuration​

WebSearch Interception Parameters​

ParameterTypeRequiredDescriptionExample
enabled_providersList[String]YesList of providers to enable web search for[openai, minimax, anthropic]
search_tool_nameStringNoSpecific search tool from search_tools config. If not set, uses first available.perplexity-search

Provider Values​

Use these values in enabled_providers:

ProviderValueProviderValue
OpenAIopenaiAnthropicanthropic
MiniMaxminimaxMistralmistral
CoherecohereFireworks AIfireworks_ai
Together AItogether_aiGroqgroq
PerplexityperplexityDeepSeekdeepseek
xAIxaiHugging Facehuggingface
OCIociVertex AIvertex_ai
BedrockbedrockAzureazure
Sagemakersagemaker_chatDatabricksdatabricks
DataRobotdatarobotVLLMhosted_vllm
HerokuherokuRAGFlowragflow
CompactifcompactifCometapicometapi
A2Aa2aBytezbytez

Search Providers​

Configure which search provider to use. LiteLLM supports multiple search providers:

Providersearch_provider ValueEnvironment Variable
Perplexity AIperplexityPERPLEXITYAI_API_KEY
TavilytavilyTAVILY_API_KEY
Exa AIexa_aiEXA_API_KEY
Brave SearchbraveBRAVE_API_KEY
Parallel AIparallel_aiPARALLEL_AI_API_KEY
Google PSEgoogle_pseGOOGLE_PSE_API_KEY, GOOGLE_PSE_ENGINE_ID
DataForSEOdataforseoDATAFORSEO_LOGIN, DATAFORSEO_PASSWORD
FirecrawlfirecrawlFIRECRAWL_API_KEY
SearXNGsearxngSEARXNG_API_BASE (required)
LinkuplinkupLINKUP_API_KEY
SerperserperSERPER_API_KEY
SearchAPI.iosearchapiSEARCHAPI_API_KEY

See Search Providers Documentation for detailed setup instructions.

Complete Configuration Example​

model_list:
# OpenAI
- model_name: gpt-5.6-terra
litellm_params:
model: openai/gpt-5.6-terra
api_key: os.environ/OPENAI_API_KEY

# MiniMax
- model_name: minimax
litellm_params:
model: minimax/MiniMax-M2.1
api_key: os.environ/MINIMAX_API_KEY

# Anthropic
- model_name: claude
litellm_params:
model: anthropic/claude-sonnet-5
api_key: os.environ/ANTHROPIC_API_KEY

# Azure OpenAI
- model_name: azure-gpt
litellm_params:
model: azure/gpt-5.6-terra
api_base: https://my-azure.openai.azure.com
api_key: os.environ/AZURE_API_KEY

litellm_settings:
callbacks: ["websearch_interception"]
websearch_interception_params:
enabled_providers:
- openai
- minimax
- anthropic
- azure
search_tool_name: perplexity-search

search_tools:
- search_tool_name: perplexity-search
litellm_params:
search_provider: perplexity
api_key: os.environ/PERPLEXITY_API_KEY

- search_tool_name: tavily-search
litellm_params:
search_provider: tavily
api_key: os.environ/TAVILY_API_KEY

- search_tool_name: parallel-search
litellm_params:
search_provider: parallel_ai
api_key: os.environ/PARALLEL_API_KEY

To use Parallel Search, set search_tool_name: parallel-search in websearch_interception_params.

Usage Examples​

Python SDK​

import litellm

# Configure callbacks
litellm.callbacks = ["websearch_interception"]

# Make completion with web search tool
response = await litellm.acompletion(
model="gpt-5.6-terra",
messages=[
{"role": "user", "content": "What are the latest AI news?"}
],
tools=[
{
"type": "function",
"function": {
"name": "litellm_web_search",
"description": "Search the web for current information",
"parameters": {
"type": "object",
"properties": {
"query": {
"type": "string",
"description": "Search query"
}
},
"required": ["query"]
}
}
}
]
)

print(response.choices[0].message.content)

Proxy Server​

# Start proxy with config
litellm --config config.yaml

# Make request
curl http://localhost:4000/v1/chat/completions \
-H "Content-Type: application/json" \
-H "Authorization: Bearer $LITELLM_API_KEY" \
-d '{
"model": "gpt-5.6-terra",
"messages": [
{"role": "user", "content": "What is the weather in San Francisco?"}
],
"tools": [
{
"type": "function",
"function": {
"name": "litellm_web_search",
"description": "Search the web",
"parameters": {
"type": "object",
"properties": {
"query": {"type": "string"}
},
"required": ["query"]
}
}
}
]
}'

How Search Tool Selection Works​

  1. If search_tool_name is specified → Uses that specific search tool
  2. If search_tool_name is not specified → Uses first search tool in search_tools list
search_tools:
- search_tool_name: perplexity-search # ← This will be used if no search_tool_name specified
litellm_params:
search_provider: perplexity
api_key: os.environ/PERPLEXITY_API_KEY

- search_tool_name: tavily-search
litellm_params:
search_provider: tavily
api_key: os.environ/TAVILY_API_KEY

Troubleshooting​

Web Search Not Working​

  1. Check provider is enabled:

    enabled_providers:
    - openai # Make sure your provider is in this list
  2. Verify search tool is configured:

    search_tools:
    - search_tool_name: perplexity-search
    litellm_params:
    search_provider: perplexity
    api_key: os.environ/PERPLEXITY_API_KEY
  3. Check API keys are set:

    export PERPLEXITY_API_KEY=your-key
  4. Enable debug logging:

    litellm.set_verbose = True

Common Issues​

Issue: Model returns tool_calls instead of final answer

  • Cause: Provider not in enabled_providers list
  • Solution: Add provider to enabled_providers

Issue: "No search tool configured" error

  • Cause: No search tools in search_tools config
  • Solution: Add at least one search tool configuration

Issue: "Invalid function arguments json string" error (MiniMax)

  • Cause: Fixed in latest version - arguments weren't properly JSON serialized
  • Solution: Update to latest LiteLLM version

Technical Details​

Architecture​

Web search integration is implemented as a custom callback (WebSearchInterceptionLogger) that:

  1. Pre-request Hook: Converts native web search tools to LiteLLM standard format
  2. Post-response Hook: Detects web search tool calls in responses
  3. Agentic Loop: Executes searches and makes follow-up requests automatically

Supported APIs​

  • ✅ Chat Completions API (OpenAI format)
  • ✅ Anthropic Messages API (Anthropic format)
  • ✅ Streaming (automatically converted)
  • ✅ Non-streaming

Response Format Detection​

The handler automatically detects response format:

  • OpenAI format: tool_calls in assistant message
  • Anthropic format: tool_use blocks in content

Performance​

  • Latency: Adds one additional LLM call (follow-up request with search results)
  • Caching: Search results can be cached (depends on search provider)
  • Parallel Searches: Multiple search queries executed in parallel

Contributing​

Found a bug or want to add support for a new provider? See our Contributing Guide.