Swarms Logo
RouteHub

One LLM gateway for every provider, ready in 3.2 milliseconds

The open-source LLM gateway behind Swarms. Call OpenAI, Anthropic, Gemini, Groq, OpenRouter, Ollama and any OpenAI-compatible server through one function, with LiteLLM's function names and 383x faster imports.

$pip install routehub

v0.2.0Apache 2.0Python 3.10+Paper

main.py
import routehub

messages = [{"role": "user", "content": "Summarize the Q3 risks."}]

for model in [
    "gpt-5.4-mini",
    "claude-sonnet-4-6",
    "gemini/gemini-3-flash-preview",
    "groq/llama-3.3-70b-versatile",
]:
    response = routehub.completion(model=model, messages=messages)
    print(response.choices[0].message.content)

Change the model string to change providers. The call and the response type stay the same.

Benchmarks

383x faster to import than LiteLLM, and 5.5x faster to a first response

Every library ran in its own environment against an instant mock server, so the numbers isolate what the client itself costs. The bare provider SDK is the floor: the same HTTP and parsing work with no gateway in front of it.

to import
3.2 ms

LiteLLM takes 1,235 ms, 383x longer

fresh process to first response
267 ms

LiteLLM takes 1,464 ms, 5.5x longer

peak memory
54 MiB

LiteLLM uses 211 MiB, 3.9x more

added to each OpenAI call
9 µs

LiteLLM adds 647 µs

Time to import the library

Median of 20 fresh processes, in milliseconds

383x faster

  • RouteHub3.2 ms
  • OpenAI SDK (bare)68x217.2 ms
  • LiteLLM386x1,235.3 ms

Importing RouteHub loads 18 modules and no third-party code. LiteLLM loads 2,502, including the OpenAI SDK, aiohttp, jinja2 and tokenizers, and with its default settings it also downloads a model price list from GitHub before your code runs.

What it adds up to over an agent run

Describe your agent and see the client-side time each gateway adds, before the model generates a single token.

50

Runs per day

4.52 s

saved on every run

13 hours of agent wall-clock time a day at 10,000 runs.

LiteLLM4.78 s
RouteHub269 ms
Where the time goesLiteLLMRouteHub
Fresh process to first response1.46 s267 ms
Gateway overhead across 50 calls37 ms1.6 ms
Reconnects after 49 tool pauses3.28 s0.0 ms

RouteHub (commit 55284e1) against LiteLLM 1.104.0 on an Apple M3 Pro with CPython 3.12, each in its own environment. The calculator uses the 22-message agent request for per-call overhead and the 67 ms handshake measured against api.openai.com; network savings scale with your round-trip time to the provider. The paper also covers async throughput, where LiteLLM's aiohttp transport pulls ahead at 128 requests in flight, and results for any-llm and aisuite.

Full methodology in the paper

Quickstart

Install it, set one key, and call any model

You need Python 3.10 or later and an API key for any supported provider. Each step builds on the one before it.

  1. 1

    Install from PyPI

    Python 3.10 or later. RouteHub has three direct dependencies: openai, pydantic and tiktoken. Add the fast extra to use orjson for JSON.

    terminal
    pip install routehub
    
    # or with uv
    uv add routehub
    
    # with orjson for faster JSON handling
    pip install "routehub[fast]"
    terminal · example output
  2. 2

    Set a key and make your first call

    Export the key for your provider, then call completion(). The response is the OpenAI SDK's own ChatCompletion, so attribute access and type hints work as usual.

    main.py
    # export OPENAI_API_KEY="sk-..."
    import routehub
    
    response = routehub.completion(
        model="gpt-5.4-mini",
        messages=[{
            "role": "user",
            "content": "Name three risks of a single-region deployment.",
        }],
    )
    print(response.choices[0].message.content)
    print(response.usage.total_tokens, "tokens")
    terminal · example output
  3. 3

    Change providers by changing one string

    Prefix the model with its provider, or use a bare name like claude-sonnet-4-6. Set the key for each provider you call. The function and the response type stay the same.

    providers.py
    import routehub
    
    messages = [{"role": "user", "content": "Reply with one word: ready?"}]
    
    for model in [
        "gpt-5.4-mini",
        "claude-sonnet-4-6",
        "gemini/gemini-3-flash-preview",
        "groq/llama-3.3-70b-versatile",
    ]:
        response = routehub.completion(model=model, messages=messages)
        print(f"{model:32} {response.choices[0].message.content}")
    terminal · example output

Features

Everything an LLM gateway needs

Pick a feature to see how it works and what it prints. The tabs above each sample switch between providers, approaches, or a before and after.

Call any model

Every major provider

OpenAI, Anthropic, Gemini, Groq, xAI, DeepSeek, OpenRouter, Together, Mistral, Fireworks, Azure OpenAI, Ollama, vLLM and any OpenAI-compatible server, all through one function.

  • Requests use the OpenAI chat format on every provider
  • Responses are the OpenAI SDK's own ChatCompletion, ChatCompletionChunk and CreateEmbeddingResponse
  • Bare names starting with gpt-, claude-, gemini-, grok-, deepseek- or mistral- need no prefix
  • api_base= points any model at a proxy, a private endpoint or a local server
hosted.py
import routehub

messages = [{"role": "user", "content": "Draft a release note."}]

routehub.completion(model="gpt-5.4", messages=messages)
routehub.completion(model="claude-sonnet-4-6", messages=messages)
routehub.completion(model="gemini/gemini-3-flash-preview", messages=messages)
routehub.completion(model="xai/grok-4", messages=messages)
routehub.completion(model="deepseek/deepseek-chat", messages=messages)
routehub.completion(
    model="openrouter/anthropic/claude-sonnet-4.6",
    messages=messages,
)

Get started

Start a new project, move off LiteLLM, or teach your coding agent

RouteHub is free and open source under Apache 2.0. Pick the path that matches where you are today.

New project

Install the package, export a key for any provider, and make your first call.

pip install routehub

Migrating from LiteLLM

Same function names and module paths. Swap the import, then move global settings into each call.

- from litellm import completion
+ from routehub import completion
- litellm.num_retries = 3
- completion(model=model, messages=messages)
+ completion(model=model, messages=messages, num_retries=3)

Coding agents

RouteHub ships an Agent Skill that shows Claude Code and other coding agents how to call it correctly, including which LiteLLM habits break.

git clone https://github.com/The-Swarm-Corporation/RouteHub.git
mkdir -p ~/.claude/skills
cp -r RouteHub/skills/routehub ~/.claude/skills/

Paper

The design and every benchmark, written up in full

FAQ

Questions about RouteHub

What Python developers ask before they switch LLM gateways.

What is RouteHub?

RouteHub is an open-source LLM gateway for Python from the Swarms team. One completion() function calls OpenAI, Anthropic, Gemini, Groq, xAI, DeepSeek, OpenRouter, Mistral, Together, Fireworks, Azure OpenAI, Ollama, vLLM and any OpenAI-compatible server, and every response comes back as the official OpenAI SDK's ChatCompletion type. Install it with pip install routehub.

Read the launch post

Is RouteHub a drop-in replacement for LiteLLM?

For most Python code, yes. RouteHub keeps LiteLLM's function names and module paths, including completion, acompletion, embedding and routehub.utils.get_model_info, so the main change is the import. Settings that LiteLLM keeps in module globals, such as retries and timeouts, become arguments to each call.

Migrate from LiteLLM step by step

How much faster is RouteHub than LiteLLM?

In the RouteHub paper's benchmarks against LiteLLM 1.104.0, RouteHub imports in 3.2 ms against 1,235 ms (383x faster), reaches a first response from a fresh process in 267 ms against 1,464 ms (5.5x faster), peaks at 54 MiB of memory against 211 MiB, and adds 9 µs to each OpenAI call against 647 µs. LiteLLM's aiohttp transport is ahead on async throughput at 128 requests in flight.

Why is LiteLLM slow?

Which LLM providers does RouteHub support?

More than 20, including OpenAI, Anthropic, Google Gemini, Groq, xAI, DeepSeek, OpenRouter, Mistral, Together, Fireworks, Azure OpenAI, Ollama and vLLM. Any server that speaks the OpenAI API works through api_base, including private endpoints and local servers. Native Bedrock and Vertex AI adapters are on the roadmap.

Does RouteHub support streaming, async, tool calling and structured output?

Yes. stream=True yields ChatCompletionChunk objects, acompletion takes the same arguments as completion for concurrent requests, tools defined once in the OpenAI format work on every provider, and passing a pydantic model returns JSON that validates against it. For Claude, RouteHub translates tool definitions, calls and results to Anthropic's format and back.

Does RouteHub support Claude prompt caching and extended thinking?

Yes. Anthropic's OpenAI-compatible endpoint drops prompt caching, thinking output and cached-token usage, so RouteHub calls Anthropic's native Messages API instead and keeps all three. In the paper's benchmarks, RouteHub's Claude path was also faster than the official Anthropic SDK for both regular and streaming calls.

Use Claude in OpenAI format from Python

Does RouteHub include a proxy server like LiteLLM's?

No. RouteHub is a Python library that runs inside your process. It has no proxy server, budgets, spend logging, callbacks or router. Teams that depend on those features can keep the LiteLLM proxy or choose a self-hosted gateway, and our comparison of LiteLLM alternatives covers both options.

Compare LiteLLM alternatives

Is RouteHub free?

Yes. RouteHub is free and open source under the Apache 2.0 license and runs on Python 3.10 and later. It has three direct dependencies (openai, pydantic and tiktoken) and installs 21 packages in total, against 58 for LiteLLM. You pay only your model providers for the tokens you use.

Compare Python LLM gateways