Product

LLM router: what it routes and what it leaves out

An LLM router picks a model for each request. It does not route the data and tool calls an agent makes. Where each kind of routing fits, and how fallback works.

Walid Boulanouar · · 7 min read

On this page

An LLM router sits in front of language models and decides which one answers each request. It might pick a cheaper model for an easy question, a stronger one for a hard question, or the next vendor when the first is down. You can run one yourself, and several open source routers exist.

An agent makes a second kind of call that a model router never sees. When the agent needs a search result, a company record or a page of text, it calls a data or tool API. Those calls fail, cost money and have several possible providers too. This post separates the two kinds of routing and explains what the second kind needs.

What is an LLM router?

An LLM router is a layer that receives a prompt and forwards it to one of several models. The decision uses rules or a classifier: the length of the prompt, the task type, a budget, or whether the preferred model is available. The caller sees one endpoint. The router handles the rest.

People want one for three reasons. Cost, because small models are cheaper and many requests do not need a large one. Reliability, because a vendor outage should not stop your product. And access, because one interface is easier than five client libraries.

Routing a model call is about text in, text out. The router never touches the world outside the model, so it cannot help with a search, a scrape or a lookup.

Is there an open source LLM router?

Yes, there are open source routers and gateways that forward model requests, and you can search for them by name. I will not rank them here, because I have not run them and a ranking without that would be a guess. If you evaluate one, test three things: how it chooses a model, what it does when a model returns an error, and whether it records the cost of each request.

Those three tests apply to any router, including the kind this post is about next.

What about the calls the model makes to tools?

An agent loop has two sides. The model reads context and decides what to do. When it decides to call a tool, your code or a gateway runs that tool and returns the result. A model router handles the first side. Something else handles the second.

The tool side has the same needs. Several providers answer the same question: Google results from one of five sources, an email lookup from one of four. Prices differ. Some providers return nothing for a given input. A good setup picks the cheapest provider that returns the fields you need and moves to the next when one fails.

Model routingTool and data routing
What it routesPrompts to language modelsCalls to search, scrape and enrichment APIs
What differs between optionsModel quality, speed and priceCoverage, fields returned and price per call
What a fallback meansTry another modelTry another provider for the same job
Who decidesA router, by ruleThe agent, with prices in hand
What you want recordedTokens and cost per requestPrice and result of each call

How does fallback work for tool calls?

In looot the catalog groups endpoints by what they do, such as "Google organic results" or "find a work email". Several providers can sit under one job. The agent finds the job with discover, reads each candidate with inspect, and sees the price and input fields before it runs anything.

Fallback is a decision the agent makes, guided by your instruction. You tell it: try the cheapest provider first, and if the result is empty or an error comes back, try the next one, and stop at a spend limit. The same idea appears in waterfall enrichment, where the agent tries email providers in order.

The live table shows how far the prices for one job can spread. That spread is the reason the order matters.

If you want to compare the SERP providers in more depth, the live SERP API ranking orders them by price.

Do I need a gateway as well as a router?

They sit in different places. A model router goes between your agent and the language models. A tool gateway goes between your agent and the data and tool APIs. You can use both, one or neither. A script with one model and two tools needs neither.

When the number of tools grows, a gateway starts to pay off, because it gives the agent one connection, one token and one record of what ran. MCP gateway for agents explains that case. MCP vs API for agents helps you choose how the agent connects.

What routing rules can I write today?

You do not need a product to start. For the tool side, a short ordered list in your agent's instruction is enough.

  1. For each job, ask for the endpoints and their prices before the first call.
  2. Run the cheapest endpoint first.
  3. If it returns an error, stop and report. If it returns an empty result, try the next cheapest endpoint once.
  4. Never try more than two providers for one item.
  5. Stop the whole task when the balance falls below the limit you named.

Rule 4 deserves a note. A fallback chain with no end turns one failed lookup into five paid calls. Two providers is a cap that keeps misses cheap. If both fail, the item is probably not findable, and a third attempt rarely helps.

How do I know whether routing saves money?

Measure the same job two ways on a small sample. Run it with the cheapest provider alone and with the fallback rule, and compare cost per found result, not cost per call. A cheaper provider that finds fewer results can cost more per result. A dearer one that finds nearly all of them can cost less.

The same logic holds for model routing. A router that sends hard questions to a cheaper model saves nothing if you re-ask the question. Count the outcomes you wanted, then divide the spend by them.

Try it: a prompt for your agent

This prompt tests fallback on a data call without spending much. It asks the agent to try the cheapest provider and then a second one for the same job, so you can see the difference.

Prompt for your agent

Test provider fallback

Test provider fallback
Use looot to compare two providers for one data job. 1. Set up https://looot.ai/skill.md. 2. Use discover to find endpoints for "google organic search results". List every match with its provider and price. Do not run anything yet. 3. Pick the cheapest and the second cheapest. Tell me their prices and wait for me to say go. 4. After I say go, run both for the query "llm router" with the same country, and compare the first five results. 5. Report the cost of each run, which result lists differ, and which provider you would try first next time.

Steps 2 and 3 are free. Step 4 costs the two prices you saw.

What it cannot do

  • It is not a model router. looot does not choose between language models or sell model inference.
  • It does not switch providers by itself. The agent follows your instruction to try the next one.
  • It does not make every provider return the same fields. Read each endpoint's output before you rely on it.
  • It does not retry on a timer. Anything that needs scheduled runs is not available yet.

Last checked 2026-10-03 against the live catalog and the looot skill file. Author: Walid.

Questions

An LLM router is a layer that forwards each request to one of several language models, using rules or a classifier. It helps with cost, reliability and a single interface to many model vendors.

ShareXLinkedInEmail
Comparisons

MCP vs API for agents: when each one fits

MCP and a plain API both let an agent reach outside data. What the agent handles in each, when to pick which, and a prompt to test both on your job.

Walid Boulanouar · · 5 min read

Try it with the agent you already use

Start your workspace, top up, and paste one prompt into your agent.