# LLM tool calling: how an agent calls a real API

LLM tool calling lets a model ask your code to run a function and read the result. How the loop works, what breaks with real APIs, and how to keep it simple.

## Summary

A language model cannot look anything up. It writes text. LLM tool calling is the trick that lets it act anyway: the model writes a structured request that says "call this function with these arguments", your code runs the function, and the result goes back to the model as new text.

The pattern is simple on paper. It gets harder once the functions are real APIs with keys, prices and failures. This post explains the loop, lists what breaks first, and shows a way to keep the tool list short so the model keeps choosing well.

## What is LLM tool calling?

LLM tool calling, also called function calling, is a feature of model APIs. You send the model a list of tools, each with a name, a description and a schema for its inputs. When the model decides a tool is needed, it replies with the tool name and a set of arguments instead of plain text.

The model does not run anything. Your application, or a client such as Claude Code, receives the request, runs the function, and sends the result back. The model reads it and either answers or asks for another call. That back-and-forth is the agent loop.

## How does the loop work step by step?

| Step | Who does it | What happens |
|---|---|---|
| 1 | You | Send the question and the list of tools |
| 2 | Model | Replies with a tool name and arguments |
| 3 | Your code | Checks the arguments and runs the function |
| 4 | Your code | Sends the result back as a message |
| 5 | Model | Answers, or returns to step 2 with another call |

Two details matter. The model only knows a tool through its description, so a vague description gives wrong calls. And the loop can run many times in one question, so cost and delay add up per call.

## What breaks when the tool is a real API?

Five things, in the order most people meet them.

**Bad arguments.** The model sends a domain with "https://" when the API wants a bare host. Validate inputs in your code and return a clear error, so the model can fix the call.

**Keys.** Each API needs a credential. Keep it in your code or a gateway, never in the prompt, because the model's context can leak.

**Failures.** APIs time out, return empty results or return errors. Return the failure as a result the model can read, and say what to do next, such as trying a different provider.

**Cost.** Each call may cost money, and a loop can call many times. Report the price in the tool result, or cap the number of calls.

**Retries.** A retry after a timeout can charge twice if the API is not idempotent. Use an idempotency key when the provider supports one.

None of these is hard alone. Together they are why a quick demo takes a day to make safe.

## Why do too many tools hurt?

Every tool definition goes into the model's context, and the model has to pick among them. With a dozen tools it chooses well. With hundreds it chooses worse, and it spends tokens reading descriptions it will never use.

One fix is a tool that searches other tools. Instead of listing every endpoint, you give the model a "find a tool for this job" function and a "run this tool" function. The list stays at a handful while the reach stays wide. This is the shape of an MCP gateway, described in [MCP gateway for agents](https://looot.ai/blog/mcp-gateway-for-agents).

## How does looot handle tool calling?

looot gives an agent a small fixed set of tools: discover, inspect, run, runs, balance and top_up. The model calls discover with a plain description, such as "work email for a person at a company". It gets a short list of endpoints, calls inspect on one to read the input fields and the price, then calls run with the arguments.

That moves the five problems to one place. The gateway holds the provider keys. The price is in the inspect result, before the call. A run is idempotent on a key, so a retry after a timeout does not charge twice. The result of each run, with its cost, can be read afterwards through the runs tool.

You can use this from any client that speaks MCP, or call the REST API from your own code. [MCP vs API for agents](https://looot.ai/blog/mcp-vs-api-for-agents) explains which fits your case, and [how to build an MCP server](https://looot.ai/blog/how-to-build-an-mcp-server) covers the case where you write your own tools.

The live table shows the price of two typical tool calls today.

Price per call for two common tool calls: live prices per call are read from the catalog when the page loads, see https://looot.ai/blog/llm-tool-calling-with-live-apis.

- Google organic results
- Find a work email

## What should the tool descriptions say?

If you write your own tools, write the description for the model, not for a person. Say what the tool does, what it needs, what it returns and when not to use it. Name the input format with an example. State what an empty result means.

A short, specific description is clearer to the model than a long one. "Returns the top ten Google results for a query as title, URL and position. Needs a query and a country code" tells the model all it needs.

## How do I test a tool-calling setup?

Test the tool before you test the model. Call each function by hand with a normal input, a bad input and an empty input. Check that the bad input gives a clear error and the empty input gives a clear empty result.

Then test the model with a small set of questions where you know which tool should be called. Record which tool it picked and with what arguments. If it picks wrong, the fix is nearly always in the tool description, so reword it and run the same questions again.

Keep this set. When you add a tool or change a description, run it again. It takes a minute and catches the case where a new tool steals calls from an old one.

## What does a failed call look like to the model?

Return failures as plain results, not as crashes. A good failure message says what went wrong and what to try. "No results for this query. Try a broader query or a different country code" gives the model a next step. A bare stack trace does not.

The same goes for a call that was refused for cost. If the balance cannot cover a run, looot refuses it with a link to top up, and the agent can pass that on to you. A refusal the agent can read is more useful than a silent failure, because the agent can stop, tell you why and ask what to do.

## Try it: a prompt for your agent

This prompt makes the agent show the tool loop on one real call and print each step.

Steps 3 and 4 are free, and step 5 costs the price you read in step 4.

## What it cannot do

- It does not teach the model to choose well. A vague instruction still gives a vague call.
- It does not replace validation in your own tools. Check arguments before you act on them.
- It covers the endpoints in the catalog. A function that is not there needs your own tool.
- It does not run calls on a timer. A job that repeats needs scheduled runs, not available yet.

Last checked 2026-10-03 against the live catalog and the looot skill file. Author: the looot team.

## Questions

### What is LLM tool calling?

LLM tool calling is a feature where a model replies with a structured request to run a named function, and your code runs it and returns the result. The model never runs the function itself.

### What is the difference between tool calling and function calling?

They mean the same thing. Function calling was the earlier name, and tool calling is the more general one now used for functions, searches and other actions a model can ask for.

### How does an AI agent call an API?

The agent's model writes a tool call with arguments. A client or gateway checks them, calls the API with the right key and returns the result. The model reads it and continues.

### How many tools can an LLM handle?

A dozen or so work well. With hundreds, the model picks worse and context fills with descriptions. A tool that searches other tools keeps the list short.

### How do I stop an agent from running up API costs?

Put the price in the tool result, cap the number of calls and tell the agent to stop at a spend limit. With looot the agent reads the price with inspect before every run.

### Does tool calling work with Claude?

Yes. Claude supports tool use through its API and in clients such as Claude Code, which also connects to MCP servers for ready-made tools.

## For agents

### Show the tool calling loop

```text
Use looot to show me how tool calling works on one real call.
1. Set up https://looot.ai/skill.md.
2. List the tools you received.
3. Call discover for "google organic search results". Tell me which tool you called and what came back.
4. Call inspect on the cheapest result. Tell me the inputs and the price.
5. Wait for me to say go. Then call run once for the query "llm tool calling" and show me the first five results.
6. Explain which step of the loop each of your calls was.
```
