Skip to main content

Shared Responsibility Model

When you self-host LiteLLM, you run the software and we build it. That split decides who debugs what. Here, we describe which problems are our responsibility and which ones are yours, so that a report reaches the right team.

In a nutshell, we own the behavior of the product as documented on this site, and you own the environment it runs in + any code you add to it.

AreaOwner
Correctness of documented features and endpointsLiteLLM
Memory leaks, hangs, and stability problems in the documented feature setLiteLLM
Provider translation, cost tracking, and routing behavior as documentedLiteLLM
Security patches and the official Docker image and Helm chartLiteLLM
Uptime of your instance and the infrastructure under itYou
Custom callbacks, custom guardrails, custom auth, and other code you injectYou
Infra issues due to deploying in a way that differs from our recommended path (your own Dockerfile, chart, or base image)You
Bugs introduced by your own patches on a fork not present in upstreamYou
Your provider accounts, quotas, and provider-side outagesYou

What we are responsible for

We are responsible for the product working. Every feature documented on this site should behave as documented. If it does not, that is a bug for us and you should open an issue or raise it in your enterprise support channel.

That responsibility covers stability, not only correctness. Memory growth, file descriptor or connection leaks, deadlocks, hangs, and throughput regressions within the documented feature set are our responsibility to diagnose and fix. This covers the interfaces you interact with: the public HTTP surface is governed by the API Stability Policy, version numbering and what a patch or minor bump means is documented in Release Cycle, and beta features moving behind Enterprise by the Migration Policy. We maintain the official Docker image, Helm chart, and Terraform modules described in Production Deployment, and we ship security patches for the supported version window.

If you are on an end-of-life line, we recommend upgrading as a first step to ensure you have the latest bug fixes and security patches applied.

What you are responsible for

You are responsible for keeping your instance up, apart from stability defects in the application itself. That means capacity and sizing, restarts and rollouts, health checking and autoscaling, and the health of Postgres, Redis, your network, and your orchestrator. Production Best Practices, Database Sizing, and Redis Sizing cover the settings and sizing we recommend. The health endpoints are there for your probes.

You are also responsible for any custom code you introduce to the gateway. Custom callbacks, custom guardrails, custom auth, custom SSO, hooks, and plugins execute in the proxy process, so a blocking call, an unbounded cache, or a leaked client in that code can show up as proxy latency, memory growth, or a hang even when the proxy is behaving correctly. The logic of your handler, and its performance and memory behavior, is under your ownership. The same applies to anything you wrap around the gateway, including sidecars, proxies in front of it, and added middleware that mutates requests.

Running a fork is the same way. A bug that also reproduces on unmodified upstream at the same version is firmly within our responsibility to debug and fix. A bug your patches introduced is under your ownership, and so is keeping those patches working as you rebase onto newer releases. If you have patched around something because upstream lacked it, create an issue or send the patch as a pull request, and if it is a general improvement, we are happy to add it upstream.

Also, if your deployment strategy is not following our recommended path, that path is yours to maintain. Plenty of teams build their own image, write their own chart, change the base image or Python version, pin their own dependency set, or run their own process manager and worker counts. That is supported use of the software. It also means a broken build, a missing system library, a mismatched dependency, an OOMKill from a container memory limit, a misconfigured worker count, etc. is something you own. See:

Filing an issue with us

Please include:

  1. The LiteLLM version
  2. How you deployed it
  3. A redacted config
  4. The exact request
  5. The full error or traceback with detailed debug logging enabled
  6. For stability reports, we recommend including the memory or latency curve over time, the request rate, and the worker and container limits
  7. For memory and latency issues, we recommend including Pyroscope profiling results

Open bugs and feature requests as GitHub issues. Enterprise customers can also use their dedicated support channel. See Professional Support for hours and SLA options.

🚅
LiteLLM Enterprise
SSO/SAML, audit logs, spend tracking, multi-team management, and guardrails — built for production.
Learn more →