LLM router: what it routes and what it leaves out
An LLM router picks a model for each request. It does not route the data and tool calls an agent makes. Where each kind of routing fits, and how fallback works.
Walid Boulanouar · · 7 min read
On this page
An LLM router sits in front of language models and decides which one answers each request. It might pick a cheaper model for an easy question, a stronger one for a hard question, or the next vendor when the first is down. You can run one yourself, and several open source routers exist.
An agent makes a second kind of call that a model router never sees. When the agent needs a search result, a company record or a page of text, it calls a data or tool API. Those calls fail, cost money and have several possible providers too. This post separates the two kinds of routing and explains what the second kind needs.
What is an LLM router?
An LLM router is a layer that receives a prompt and forwards it to one of several models. The decision uses rules or a classifier: the length of the prompt, the task type, a budget, or whether the preferred model is available. The caller sees one endpoint. The router handles the rest.
People want one for three reasons. Cost, because small models are cheaper and many requests do not need a large one. Reliability, because a vendor outage should not stop your product. And access, because one interface is easier than five client libraries.
Routing a model call is about text in, text out. The router never touches the world outside the model, so it cannot help with a search, a scrape or a lookup.
Is there an open source LLM router?
Yes, there are open source routers and gateways that forward model requests, and you can search for them by name. I will not rank them here, because I have not run them and a ranking without that would be a guess. If you evaluate one, test three things: how it chooses a model, what it does when a model returns an error, and whether it records the cost of each request.
Those three tests apply to any router, including the kind this post is about next.
What about the calls the model makes to tools?
An agent loop has two sides. The model reads context and decides what to do. When it decides to call a tool, your code or a gateway runs that tool and returns the result. A model router handles the first side. Something else handles the second.
The tool side has the same needs. Several providers answer the same question: Google results from one of five sources, an email lookup from one of four. Prices differ. Some providers return nothing for a given input. A good setup picks the cheapest provider that returns the fields you need and moves to the next when one fails.
| Model routing | Tool and data routing | |
|---|---|---|
| What it routes | Prompts to language models | Calls to search, scrape and enrichment APIs |
| What differs between options | Model quality, speed and price | Coverage, fields returned and price per call |
| What a fallback means | Try another model | Try another provider for the same job |
| Who decides | A router, by rule | The agent, with prices in hand |
| What you want recorded | Tokens and cost per request | Price and result of each call |
How does fallback work for tool calls?
In looot the catalog groups endpoints by what they do, such as "Google organic results" or "find a work email". Several providers can sit under one job. The agent finds the job with discover, reads each candidate with inspect, and sees the price and input fields before it runs anything.
Fallback is a decision the agent makes, guided by your instruction. You tell it: try the cheapest provider first, and if the result is empty or an error comes back, try the next one, and stop at a spend limit. The same idea appears in waterfall enrichment, where the agent tries email providers in order.
The live table shows how far the prices for one job can spread. That spread is the reason the order matters.
If you want to compare the SERP providers in more depth, the live SERP API ranking orders them by price.
Do I need a gateway as well as a router?
They sit in different places. A model router goes between your agent and the language models. A tool gateway goes between your agent and the data and tool APIs. You can use both, one or neither. A script with one model and two tools needs neither.
When the number of tools grows, a gateway starts to pay off, because it gives the agent one connection, one token and one record of what ran. MCP gateway for agents explains that case. MCP vs API for agents helps you choose how the agent connects.
What routing rules can I write today?
You do not need a product to start. For the tool side, a short ordered list in your agent's instruction is enough.
- For each job, ask for the endpoints and their prices before the first call.
- Run the cheapest endpoint first.
- If it returns an error, stop and report. If it returns an empty result, try the next cheapest endpoint once.
- Never try more than two providers for one item.
- Stop the whole task when the balance falls below the limit you named.
Rule 4 deserves a note. A fallback chain with no end turns one failed lookup into five paid calls. Two providers is a cap that keeps misses cheap. If both fail, the item is probably not findable, and a third attempt rarely helps.
How do I know whether routing saves money?
Measure the same job two ways on a small sample. Run it with the cheapest provider alone and with the fallback rule, and compare cost per found result, not cost per call. A cheaper provider that finds fewer results can cost more per result. A dearer one that finds nearly all of them can cost less.
The same logic holds for model routing. A router that sends hard questions to a cheaper model saves nothing if you re-ask the question. Count the outcomes you wanted, then divide the spend by them.
Try it: a prompt for your agent
This prompt tests fallback on a data call without spending much. It asks the agent to try the cheapest provider and then a second one for the same job, so you can see the difference.
Prompt for your agent
Test provider fallback
Steps 2 and 3 are free. Step 4 costs the two prices you saw.
What it cannot do
- It is not a model router. looot does not choose between language models or sell model inference.
- It does not switch providers by itself. The agent follows your instruction to try the next one.
- It does not make every provider return the same fields. Read each endpoint's output before you rely on it.
- It does not retry on a timer. Anything that needs scheduled runs is not available yet.
Last checked 2026-10-03 against the live catalog and the looot skill file. Author: Walid.
Questions
Keep reading
MCP gateway: what it does and when an agent needs one
An MCP gateway puts many tools behind one connection, one login and one bill. How it works, when you need one, and what looot's gateway does.
The looot team · · 7 min read
MCP vs API for agents: when each one fits
MCP and a plain API both let an agent reach outside data. What the agent handles in each, when to pick which, and a prompt to test both on your job.
Walid Boulanouar · · 5 min read
Waterfall enrichment explained: how it works and where it breaks
Waterfall enrichment tries one data provider, then the next on a miss. How it works, why order matters, a prompt to run it with an agent, and the limits.
Walid Boulanouar · · 5 min read
Try it with the agent you already use
Start your workspace, top up, and paste one prompt into your agent.