LiteLLM: Call 100+ LLM APIs With One Unified OpenAI Interface
LiteLLM normalizes the APIs of OpenAI, Anthropic, Bedrock, Gemini, and 100+ other providers into one OpenAI-compatible interface with built-in fallbacks and cost tracking.
Mahmudul Haque Qudrati
CEO & ML Engineer
One AI engineering post, weekly
LLM benchmarks, prompt techniques, and token-cost breakdowns — not another AI news roundup.
If you write code against the OpenAI SDK and later want to switch to Claude or Gemini, you rewrite your API calls. If you want fallback logic (try GPT-4o, fall back to Claude if it's down), you implement it yourself. If you want to track spend across providers, you build that too.
LiteLLM solves all three problems with a single library and an optional proxy server.
The completion() Function
from litellm import completion
# OpenAI
response = completion(model="gpt-4o", messages=[{"role": "user", "content": "Hello"}])
# Anthropic — same function, same response format
response = completion(model="claude-3-5-sonnet-20241022", messages=[{"role": "user", "content": "Hello"}])
# Bedrock Claude
response = completion(model="bedrock/anthropic.claude-3-5-sonnet-20241022-v2:0", messages=[{"role": "user", "content": "Hello"}])
# Gemini
response = completion(model="gemini/gemini-1.5-pro", messages=[{"role": "user", "content": "Hello"}])
Every response comes back in OpenAI format regardless of provider.
Team workspace
Ship faster with chat, meetings, and projects in one place — Zlyqor.
Fallback Logic
from litellm import completion
response = completion(
model="gpt-4o",
messages=[{"role": "user", "content": "Hello"}],
fallbacks=["claude-3-5-sonnet-20241022", "gemini/gemini-1.5-pro"],
num_retries=2,
)
If GPT-4o returns an error or rate limit, LiteLLM automatically tries Claude, then Gemini.
The Proxy Server
LiteLLM's proxy turns any model into an OpenAI-compatible endpoint. Any tool that accepts an OpenAI API URL can point to your LiteLLM proxy:
litellm --model claude-3-5-sonnet-20241022 --port 8000
Then use the standard OpenAI SDK pointed at localhost:
from openai import OpenAI
client = OpenAI(base_url="http://localhost:8000", api_key="anything")
response = client.chat.completions.create(
model="claude-3-5-sonnet-20241022",
messages=[{"role": "user", "content": "Hello"}],
)
Cost Tracking and Virtual Keys
The proxy includes built-in cost tracking per request, per key, and per team. Virtual API keys let you issue per-team keys with spend limits:
# config.yaml for litellm proxy
model_list:
- model_name: gpt-4o
litellm_params:
model: gpt-4o
api_key: os.environ/OPENAI_API_KEY
- model_name: claude-sonnet
litellm_params:
model: claude-3-5-sonnet-20241022
api_key: os.environ/ANTHROPIC_API_KEY
litellm_settings:
success_callback: ["langfuse"]
budget_manager: True
Router for Load Balancing and A/B Testing
from litellm import Router
router = Router(
model_list=[
{"model_name": "fast", "litellm_params": {"model": "gpt-4o-mini"}},
{"model_name": "smart", "litellm_params": {"model": "gpt-4o"}},
],
routing_strategy="latency-based-routing",
)
response = router.completion(model="fast", messages=[...])
Resources

Mahmudul Haque Qudrati
CEO & ML Engineer
Visionary technologist, software engineer, and machine learning specialist. Founder and CEO of Pristren, directing engineering teams that ship production-grade AI/ML pipelines, mission-critical full-stack applications, and developer tooling. Creator of Zlyqor, the unified team workspace platform. Author of 540+ technical guides and benchmark research reports on large language models, agentic workflows, Model Context Protocol (MCP), and modern web stacks.
More from Mahmudul
Related Articles
How to Use Claude to Make Videos Like Vox and Others
Claude can help you make Vox-style videos by generating scripts, editing with code, and automating animation. Here's a practical guide with real workflows and costs.
OpenAI Ends Cursor Model Access on Nov 12, 2026: What Developers Need to Know
OpenAI will terminate Cursor's access to its models on November 12, 2026, following SpaceX's acquisition. This guide explains the timeline, why it happened, and practical steps to migrate your workflow.
I Used Claude Code to Get a Second Opinion on My MRI: A Practical Overview
A developer fed his MRI scan to Claude Code and got a second opinion. Here's how the experiment worked, what it cost, and why you shouldn't rely on it for medical decisions.
// discussion
Comments