Overview
LiteLLM is an open-source library for calling more than 100 large language models through a unified interface. It supports providers such as OpenAI, Anthropic, Vertex AI, Bedrock, Ollama, and Azure OpenAI while using the OpenAI request and response format.
Developers can use LiteLLM directly as a Python SDK or deploy its self-hosted AI Gateway. This makes it possible to change providers without learning a separate completion interface for each one.
Key Features
- Consistent completion calls and output structures across supported providers
- Non-streaming model responses and streaming response chunks
- Built-in retries and fallbacks across deployments through the Router
- Virtual keys, cost tracking, and an administrative interface in the gateway
- Support for function calling, prompt caching, and model routing workflows
How It Works
The Python SDK accepts a provider-qualified model name and a list of messages through the same completion interface. Authentication and provider-specific settings are configured for the selected backend, while the returned data follows the OpenAI Chat Completions structure.
For centralized deployment, the gateway can run in a container without requiring a local Python setup. Applications then send requests through one gateway rather than integrating separately with every model provider.
Agent Gateway
LiteLLM also supports A2A agents through its Agent Gateway. Agents can be registered in the administrative interface or declared in gateway configuration, then invoked with A2A protocol version 0.3 or 1.0.
Supported agent backends include A2A-compatible services, Vertex AI Agent Engine, LangGraph, Azure AI Foundry, Bedrock AgentCore, and Pydantic AI. The gateway adds request and response logging, load balancing, streaming, iteration budgets, and access controls that determine which teams or keys may use an agent.
Routing Options
The optional Auto Router gives clients one model name while classifying requests and selecting a configured model tier. A tier can contain a single model or a pool, and available classifier approaches include heuristics, an LLM, keyword rules, JEV, and custom plugins. It also supports context-window escalation, modality routing, prompt caching, and optional session pinning.
