Skip to main content

Reka

Overview​

PropertyDetails
DescriptionReka serves its own models and a curated selection of open models over one OpenAI-compatible API, with automatic prompt caching and no platform fee or markup.
Provider Route on LiteLLMreka/
Link to Provider DocReka Developer Reference ↗
Base URLhttps://api.reka.ai/v1
Supported Operations/chat/completions, /responses, /messages


We support ALL Reka models, just set reka/ as a prefix when sending requests

Available Models​

ModelContextMax output
reka/reka-flash-364k58,982
reka/reka-edge-260316k14,745
reka/deepseek4-flash1M384,000
reka/deepseek-v4-pro1M393,216
reka/glm5.3262k131,072
reka/glm5.3-flash262k131,072
reka/qwen3.8-27b262k131,072

Required Variables​

Environment Variables
os.environ["REKA_API_KEY"] = ""  # your Reka API key
os.environ["REKA_API_BASE"] = "" # optional, defaults to https://api.reka.ai/v1

Usage - LiteLLM Python SDK​

Non-streaming​

Reka Non-streaming Completion
import os
from litellm import completion

os.environ["REKA_API_KEY"] = "" # your Reka API key

response = completion(
model="reka/reka-flash-3",
messages=[{"role": "user", "content": "Hello, how are you?"}],
)

print(response.choices[0].message.content)

Streaming​

Reka Streaming Completion
import os
from litellm import completion

os.environ["REKA_API_KEY"] = "" # your Reka API key

response = completion(
model="reka/reka-flash-3",
messages=[{"role": "user", "content": "Write a short story about AI"}],
stream=True,
)

for chunk in response:
print(chunk)

Vision​

reka-edge-2603 accepts images and video. Image input uses the OpenAI content-part shape; video uses a video_url part the same way. Check input_modalities on GET /v1/models before sending media to any other model.

Reka Image Input
import os
from litellm import completion

os.environ["REKA_API_KEY"] = "" # your Reka API key

response = completion(
model="reka/reka-edge-2603",
messages=[{
"role": "user",
"content": [
{"type": "text", "text": "What animal is this? Answer briefly."},
{"type": "image_url", "image_url": {"url": "https://v0.docs.reka.ai/_images/000000245576.jpg"}},
],
}],
)

print(response.choices[0].message.content)

Responses API​

Reka does not serve /v1/responses natively, so LiteLLM translates litellm.responses calls into Reka chat completions and converts the result back into a Responses API object.

Reka Responses API
import os
import litellm

os.environ["REKA_API_KEY"] = "" # your Reka API key

response = litellm.responses(
model="reka/reka-flash-3",
input="Say hello",
)

print(response.output_text)

Usage - LiteLLM Proxy Server​

config.yaml
model_list:
- model_name: reka-flash-3
litellm_params:
model: reka/reka-flash-3
api_key: os.environ/REKA_API_KEY
- model_name: reka-edge
litellm_params:
model: reka/reka-edge-2603
api_key: os.environ/REKA_API_KEY

general_settings:
master_key: os.environ/LITELLM_MASTER_KEY

Start the proxy:

Start LiteLLM Proxy
export REKA_API_KEY="your-api-key"
export LITELLM_MASTER_KEY="sk-local-reka"
litellm --config config.yaml --port 4000

# RUNNING on http://0.0.0.0:4000

A deployment configured this way serves /v1/chat/completions, /v1/responses, and /v1/messages on the proxy.

curl
curl http://localhost:4000/v1/chat/completions \
-H "Authorization: Bearer $LITELLM_MASTER_KEY" \
-H "Content-Type: application/json" \
-d '{
"model": "reka-flash-3",
"messages": [{"role": "user", "content": "Hello, how are you?"}]
}'

You can also add Reka from the Admin UI. Go to Models, then Add Model, pick Reka as the provider, enter a reka/ model id, and paste your key.

Cost Tracking​

Reka models are not yet in LiteLLM's model cost map, so spend is not computed automatically. Reka bills per token at each model's rate with no platform fee, and publishes the rates on developer.reka.ai/models and in the pricing object of GET /v1/models (US dollars per token, as strings: prompt, completion, and input_cache_read). Pass those values as input_cost_per_token and output_cost_per_token on the deployment and LiteLLM will track spend for it, returning the amount in the x-litellm-response-cost response header and recording it in spend logs. Reka lists rates per million tokens; divide by 1,000,000 for the per-token value.

config.yaml
model_list:
- model_name: reka-edge
litellm_params:
model: reka/reka-edge-2603
api_key: os.environ/REKA_API_KEY
input_cost_per_token: 0.0000001 # $0.10 / 1M
output_cost_per_token: 0.0000001 # $0.10 / 1M
- model_name: glm5.3-flash
litellm_params:
model: reka/glm5.3-flash
api_key: os.environ/REKA_API_KEY
input_cost_per_token: 0.00000015 # $0.15 / 1M
output_cost_per_token: 0.0000005 # $0.50 / 1M

Where a model supports prompt caching, Reka caches repeated prompt prefixes automatically and bills those tokens at the cached-input rate; usage.reasoning_tokens is included inside completion_tokens and is not billed twice.

Custom API Base​

Option 1: Environment variable

Custom API Base via env var
import os
from litellm import completion

os.environ["REKA_API_BASE"] = "https://custom.reka.example/v1"
os.environ["REKA_API_KEY"] = "" # your API key

response = completion(
model="reka/reka-flash-3",
messages=[{"role": "user", "content": "Hello!"}],
)

Option 2: Pass directly

Custom API Base via parameter
from litellm import completion

response = completion(
model="reka/reka-flash-3",
messages=[{"role": "user", "content": "Hello!"}],
api_base="https://custom.reka.example/v1",
api_key="your-api-key",
)

Passing api_base="https://api.reka.ai/v1" without the reka/ prefix also resolves to the Reka provider, so model="reka-flash-3" with that base URL is routed as reka and picks up REKA_API_KEY.