Introducing litellm-core: the first step toward a leaner SDK
We're introducing litellm-core as the first step in ongoing work to make the LiteLLM SDK leaner. We want applications to install fewer dependencies, use less disk space, and spend less time loading the SDK
This first step delivers an approximately 65.6 MB reduction in installed disk space, including runtime dependencies, compared with baseline litellm in our source-build comparison. We'll continue building on that progress with further dependency reductions and improvements to how the SDK loads
The first step separates SDK packaging from the existing distribution, removes mandatory AWS and Python Hugging Face tokenizer dependencies from core, and leaves out the bundled dashboard. Your application continues to use import litellm
Why a separate core package
LiteLLM supports a broad range of providers, endpoints, and integrations. An application making ordinary chat completion calls may not need AWS credential handling or Python Hugging Face tokenizers, but those packages have been part of the SDK's default installation
Each mandatory dependency contributes to the installation footprint, and it can bring its own dependencies. Importing modules before they are needed can also add startup work. We want developers to install and load less when their application needs less
litellm-core gives us a dedicated distribution to improve that experience while preserving the existing litellm package's dependency defaults
The same SDK APIs
Core is built from the shared LiteLLM SDK sources. It retains provider request and response translation, synchronous and asynchronous completions, streaming, embeddings, token counting, model metadata, and cost calculations. Routing, retries, fallbacks, and consistent provider error handling remain part of the SDK
The broader API surface, including Responses, image, audio, and batch operations, is preserved too. Support still depends on the model and provider, and features with optional dependencies need those packages installed
The package name changes to litellm-core, but Python imports continue to use litellm. Core does not bundle the Admin UI or expose the gateway CLI entry points
What changes in this first step
Core no longer requires boto3, tokenizers, or huggingface-hub in every installation. Applications that need AWS credential resolution, request signing, or event-stream decoding can install boto3. Applications using Python Hugging Face tokenizer paths can install tokenizers and huggingface-hub
The shared implementation loads these optional packages at the operations that need them and provides installation guidance when they are missing. The proxy CLI is also deferred rather than imported as part of SDK initialization
Token counting remains available through the supported native and tiktoken paths. If a token-counting path falls back to an approximation, it emits a warning so applications can distinguish an estimate from a model-specific count
Core also declares python-dateutil directly. Removing three mandatory packages does not mean exactly three fewer installed packages: the total depends on their transitive dependencies and the rest of the application's environment
Approximately 65.6 MB less installed disk space
Our source-build comparison measured an approximately 65.6 MB reduction in installed disk space, including runtime dependencies. This is the first measured step in our ongoing work to reduce what applications need to install
The comparison covers default installations without optional extras. Core retains the shared SDK APIs, excludes bundled dashboard assets and gateway CLI entry points, and makes AWS and Python Hugging Face tokenizer packages optional. Adding those packages changes the installation footprint
How we measured
- We built baseline
litellmat commit 5d207d85 andlitellm-coreat commit 029a626c, then installed their default dependencies in separate clean environments on the same Ubuntu x86_64 machine and CPython version, with matching versions of shared dependencies. These measurements describe those source builds - Installed size counts unique files belonging to the SDK and its runtime dependencies, including metadata and scripts. It excludes the interpreter, common tooling, and generated bytecode. MB uses decimal units, with 1 MB equal to 1,000,000 bytes. The reduction measures disk space, not RAM usage or inference speed
A shared foundation for LiteLLM
Our longer-term direction is for litellm-core to become the common SDK foundation for LiteLLM, with litellm eventually wrapping it
For developers familiar with Pydantic's architecture, the separation between pydantic and pydantic-core is a useful analogy: shared core functionality can evolve behind a familiar public interface
That analogy describes the architectural direction. In this first step, litellm and litellm-core are independently packaged from shared sources. litellm has not yet become a wrapper around core
Getting started
In a fresh virtual environment, install litellm-core and continue using your existing litellm imports. Add the optional AWS or Python tokenizer packages if your application uses those features
Choose one distribution per environment. litellm and litellm-core currently install overlapping Python files and cannot coexist, even at matching versions. Existing applications can continue using litellm with its current dependency defaults
The LiteLLM Core guide covers installation, dependency choices, and the gateway distinction
Continuing to make core leaner and faster
Our longer-term goal is to bring the default installation to around 10 total dependencies, including transitive dependencies. That is a target for ongoing work, not the dependency count of this first release
We'll continue shrinking the installation footprint and loading runtime dependencies only when the relevant functionality needs them
Faster imports remain a goal of that work, alongside fewer dependencies and a smaller installation. We'll keep measuring progress as we improve how the SDK loads and uses its runtime dependencies
As core becomes the common SDK foundation and litellm eventually wraps it, we'll keep measuring these changes so developers can follow the progress
