Skip to main content

VoyageAI by MongoDB

https://docs.voyageai.com/embeddings/

Voyage AI is now VoyageAI by MongoDB. LiteLLM keeps the voyage/ model prefix and picks the API host from the key you give it, the same way the official voyageai SDK does

API Key​

# env variable
os.environ['VOYAGE_API_KEY']

LiteLLM also reads VOYAGE_AI_API_KEY and VOYAGE_AI_TOKEN when VOYAGE_API_KEY is not set

Which host your key goes to​

Key issued byKey prefixRequests go to
MongoDB Atlas (Model API)al-https://ai.mongodb.com/v1
Voyage AI dashboardpa-https://api.voyageai.com/v1

Each host rejects the other's keys with a 403, so there is nothing to configure: a MongoDB-issued key is routed to ai.mongodb.com on its own. Set api_base on the model (or VOYAGE_API_BASE for rerank) to send requests somewhere else, for example a gateway in front of either host; an explicit api_base always wins over the key prefix

model_list:
- model_name: voyage-3.5
litellm_params:
model: voyage/voyage-3.5
api_key: os.environ/VOYAGE_API_KEY # al-... goes to ai.mongodb.com, pa-... to api.voyageai.com
- model_name: rerank-2.5
litellm_params:
model: voyage/rerank-2.5
api_key: os.environ/VOYAGE_API_KEY

Sample Usage - Embedding​

from litellm import embedding
import os

os.environ['VOYAGE_API_KEY'] = ""
response = embedding(
model="voyage/voyage-3.5",
input=["good morning from litellm"],
)
print(response)

Supported Parameters​

VoyageAI embeddings support the following optional parameters:

  • input_type: Specifies the type of input for retrieval optimization
    • "query": Use for search queries
    • "document": Use for documents being indexed
  • dimensions: Output embedding dimensions (256, 512, 1024, or 2048)
  • encoding_format: Output format ("float", "int8", "uint8", "binary", "ubinary")
  • truncation: Whether to truncate inputs exceeding max tokens (default: True)

Example with Parameters​

from litellm import embedding
import os

os.environ['VOYAGE_API_KEY'] = "your-api-key"

# Embedding with custom dimensions and input type
response = embedding(
model="voyage/voyage-3.5",
input=["Your text here"],
dimensions=512,
input_type="document"
)
print(f"Embedding dimensions: {len(response.data[0]['embedding'])}")

Supported Models​

All models listed here https://docs.voyageai.com/embeddings/#models-and-specifics are supported

Model NameFunction Call
voyage-4-largeembedding(model="voyage/voyage-4-large", input)
voyage-4embedding(model="voyage/voyage-4", input)
voyage-4-liteembedding(model="voyage/voyage-4-lite", input)
voyage-code-4embedding(model="voyage/voyage-code-4", input)
voyage-context-4embedding(model="voyage/voyage-context-4", input)
voyage-context-3embedding(model="voyage/voyage-context-3", input)
voyage-3.5embedding(model="voyage/voyage-3.5", input)
voyage-3.5-liteembedding(model="voyage/voyage-3.5-lite", input)
voyage-3-largeembedding(model="voyage/voyage-3-large", input)
voyage-3embedding(model="voyage/voyage-3", input)
voyage-3-liteembedding(model="voyage/voyage-3-lite", input)
voyage-code-3embedding(model="voyage/voyage-code-3", input)
voyage-finance-2embedding(model="voyage/voyage-finance-2", input)
voyage-law-2embedding(model="voyage/voyage-law-2", input)
voyage-code-2embedding(model="voyage/voyage-code-2", input)
voyage-multilingual-2embedding(model="voyage/voyage-multilingual-2", input)
voyage-large-2-instructembedding(model="voyage/voyage-large-2-instruct", input)
voyage-large-2embedding(model="voyage/voyage-large-2", input)
voyage-2embedding(model="voyage/voyage-2", input)
voyage-lite-02-instructembedding(model="voyage/voyage-lite-02-instruct", input)
voyage-01embedding(model="voyage/voyage-01", input)
voyage-lite-01embedding(model="voyage/voyage-lite-01", input)
voyage-lite-01-instructembedding(model="voyage/voyage-lite-01-instruct", input)

Contextual Embeddings (voyage-context-4, voyage-context-3)​

Voyage's voyage-context-4 and voyage-context-3 models produce contextualized chunk embeddings: each chunk is embedded with awareness of the whole document it came from, which retrieves better on long documents than embedding the chunks on their own. LiteLLM sends any Voyage model with context in its name to Voyage's /v1/contextualizedembeddings endpoint, so the same embedding() call and /v1/embeddings proxy route work; only the input and response shapes differ from the regular models

Input shapes​

A flat list of strings, or a single string, embeds each string as its own document. LiteLLM forwards it with enable_auto_chunking: true, chunk_size: 32000, and input_type: "document", so a string of up to 32,000 tokens comes back as one embedding and a longer one is split into chunks of up to 32,000 tokens on Voyage's side. Sending input_type: "query" skips those defaults and embeds each string as a search query. Any input_type, chunk_size, or enable_auto_chunking you pass yourself replaces the default

from litellm import embedding
import os

os.environ['VOYAGE_API_KEY'] = "your-api-key"

# Each string is embedded as its own document
response = embedding(
model="voyage/voyage-context-4",
input=["The quick brown fox", "jumps over the lazy dog"],
)
print(f"Documents embedded: {len(response.data)}")

# Search queries
response = embedding(
model="voyage/voyage-context-4",
input=["what does the fox do", "who is lazy"],
input_type="query",
)

A nested list is the pre-chunked form: each inner list is one document you already split into chunks, and LiteLLM forwards it unchanged

# Single document with multiple chunks
response = embedding(
model="voyage/voyage-context-4",
input=[
[
"Chapter 1: Introduction to AI",
"This chapter covers the basics of artificial intelligence.",
"We will explore machine learning and deep learning."
]
]
)
print(f"Number of chunk groups: {len(response.data)}")

# Multiple documents
response = embedding(
model="voyage/voyage-context-4",
input=[
["Paris is the capital of France.", "It is known for the Eiffel Tower."],
["Tokyo is the capital of Japan.", "It is a major economic hub."]
]
)
print(f"Processed {len(response.data)} documents")

Response shape​

The response keeps Voyage's nested layout: data has one entry per input, and that entry's data holds one embedding per chunk. response.data[0]["data"][0]["embedding"] is the first chunk of the first input, which with flat input and the default chunk size is the whole string

LiteLLM Proxy​

Add the model to config.yaml:

model_list:
- model_name: voyage-context-4
litellm_params:
model: voyage/voyage-context-4
api_key: os.environ/VOYAGE_API_KEY

Flat list, one document per string:

curl http://localhost:4000/v1/embeddings \
-H "Authorization: Bearer $LITELLM_API_KEY" \
-H "Content-Type: application/json" \
-d '{
"model": "voyage-context-4",
"input": ["The quick brown fox", "jumps over the lazy dog"]
}'

Flat list as search queries:

curl http://localhost:4000/v1/embeddings \
-H "Authorization: Bearer $LITELLM_API_KEY" \
-H "Content-Type: application/json" \
-d '{
"model": "voyage-context-4",
"input": ["what does the fox do", "who is lazy"],
"input_type": "query"
}'

Nested list, one document already split into chunks:

curl http://localhost:4000/v1/embeddings \
-H "Authorization: Bearer $LITELLM_API_KEY" \
-H "Content-Type: application/json" \
-d '{
"model": "voyage-context-4",
"input": [["The quick brown fox", "jumps over the lazy dog"]]
}'

Specifications​

ModelPer chunkPer requestOutput dimensionsPrice/M tokens
voyage-context-432,000 tokens120,000 tokens, 1,000 inputs, 16,000 chunks256, 512, 1024 (default), 2048$0.12
voyage-context-332,000 tokens120,000 tokens, 1,000 inputs, 16,000 chunks256, 512, 1024 (default), 2048$0.18

The limits are Voyage's, from https://docs.voyageai.com/docs/contextualized-chunk-embeddings, and the per-request token total counts every chunk in the call

When to use contextual embeddings​

Reach for voyage-context-4 when you split long documents into chunks and the surrounding document should inform each chunk's embedding, because structure, section references, and cross-chunk dependencies matter. Reach for voyage-4-large, voyage-4, or voyage-4-lite for independent pieces of text and short queries, where document context adds nothing and the standard models are cheaper and faster

Model Selection Guide​

ModelBest ForContext LengthPrice/M Tokens
voyage-4-largeBest general-purpose and multilingual quality32K$0.12
voyage-4General-purpose, multilingual32K$0.06
voyage-4-liteLatency-sensitive applications32K$0.02
voyage-code-4Code retrieval and coding agents32K$0.12
voyage-context-4Contextual document embeddings32K per chunk, 120K per request$0.12
voyage-3.5General-purpose, multilingual32K$0.06
voyage-3.5-liteLatency-sensitive applications32K$0.02
voyage-3-largeBest overall quality32K$0.18
voyage-code-3Code retrieval and search32K$0.18
voyage-finance-2Financial documents32K$0.12
voyage-law-2Legal documents16K$0.12
voyage-context-3Contextual document embeddings32K per chunk, 120K per request$0.18

Rerank​

VoyageAI by MongoDB provides reranking models to improve search relevance by reordering documents based on their relevance to a query. A MongoDB-issued key (al-) is routed to ai.mongodb.com here too.

Quick Start​

from litellm import rerank
import os

os.environ["VOYAGE_API_KEY"] = "your-api-key"

response = rerank(
model="voyage/rerank-2.5",
query="What is the capital of France?",
documents=[
"Paris is the capital of France.",
"London is the capital of England.",
"Berlin is the capital of Germany.",
],
top_n=3,
)

print(response)

Async Usage​

from litellm import arerank
import os
import asyncio

os.environ["VOYAGE_API_KEY"] = "your-api-key"

async def main():
response = await arerank(
model="voyage/rerank-2.5-lite",
query="Best programming language for beginners?",
documents=[
"Python is great for beginners due to simple syntax.",
"JavaScript runs in browsers and is versatile.",
"Rust has a steep learning curve but is very safe.",
],
top_n=2,
)
print(response)

asyncio.run(main())

LiteLLM Proxy Usage​

Add to your config.yaml:

model_list:
- model_name: rerank-2.5
litellm_params:
model: voyage/rerank-2.5
api_key: os.environ/VOYAGE_API_KEY
- model_name: rerank-2.5-lite
litellm_params:
model: voyage/rerank-2.5-lite
api_key: os.environ/VOYAGE_API_KEY

Test with curl:

curl http://localhost:4000/rerank \
-H "Authorization: Bearer $LITELLM_API_KEY" \
-H "Content-Type: application/json" \
-d '{
"model": "rerank-2.5",
"query": "What is the capital of France?",
"documents": [
"Paris is the capital of France.",
"London is the capital of England.",
"Berlin is the capital of Germany."
],
"top_n": 3
}'

Supported Rerank Models​

ModelContext LengthDescriptionPrice/M Tokens
rerank-2.532KBest quality, multilingual, instruction-following$0.05
rerank-2.5-lite32KOptimized for latency and cost$0.02
rerank-216KLegacy model$0.05
rerank-2-lite8KLegacy model, faster$0.02

Supported Parameters​

ParameterTypeDescription
modelstringModel name (e.g., voyage/rerank-2.5)
querystringThe search query
documentslistList of documents to rerank
top_nintNumber of top results to return
return_documentsboolWhether to include document text in response