---
title: "Harbor"
url: "/docs/projects/Harbor"
canonical_url: "https://docs.litellm.ai/docs/projects/Harbor"
type: "docs"
last_updated: "2026-10-04"
summary: "Harbor is a framework from the creators of Terminal-Bench for evaluating and optimizing agents and language models. It uses LiteLLM to call 100+ LLM providers."
related:
  - "/docs/projects/Agent Lightning"
  - "/docs/projects/CompatCanary"
---
# Harbor

> Index of all LiteLLM docs: https://docs.litellm.ai/llms.txt


[Harbor](https://github.com/laude-institute/harbor) is a framework from the creators of Terminal-Bench for evaluating and optimizing agents and language models. It uses LiteLLM to call 100+ LLM providers.

```bash
# Install
uv add harbor

# Run a benchmark with any LiteLLM-supported model
harbor run --dataset terminal-bench@2.0 \
   --agent claude-code \
   --model anthropic/claude-opus-5 \
   --n-concurrent 4
```

Key features:
- Evaluate agents like Claude Code, OpenHands, Codex CLI
- Build and share benchmarks and environments
- Run experiments in parallel across cloud providers (Daytona, Modal)
- Generate rollouts for RL optimization

- [GitHub](https://github.com/laude-institute/harbor)
- [Documentation](https://harborframework.com/docs)

## Related pages

- [Agent Lightning](https://docs.litellm.ai/docs/projects/Agent Lightning.md)
- [CompatCanary](https://docs.litellm.ai/docs/projects/CompatCanary.md)
