# AI agent cost: model tokens plus data calls

AI agent cost has two meters, model tokens and data calls. How to estimate each, why the data side is easier to cap, and how to read per-call prices first.

## Summary

Ask what an AI agent costs and you get two answers, because an agent runs two meters. One counts the tokens the model reads and writes. The other counts the calls the agent makes to outside services: searches, scrapes, lookups.

People watch the first meter and forget the second until a bill arrives. This post explains both, shows how to estimate each for a job, and describes a way to cap the data side before the agent starts.

## What makes up AI agent cost?

AI agent cost is the sum of model cost and tool cost. Model cost is the price of the tokens the model processes. Tool cost is the price of each call the agent makes through its tools. A third part, hosting, applies if you run the agent yourself.

Each meter behaves differently. Model cost grows with the length of the conversation, since every step re-reads the earlier text. Tool cost grows with the number of calls, and each call has its own price. A long agent run with many search calls pays on both meters at once.

For model prices, read your model vendor's price page. I do not quote them here, because they change and they are not looot's to state. The rest of this post is about the tool meter, which you can read from the catalog.

## How do I estimate the data cost of a job?

Multiply three numbers. The number of items, the number of calls per item and the price of each call.

| Input | Where it comes from | Example question |
|---|---|---|
| Items | Your list or task | How many companies? |
| Calls per item | Your design | Search, then page read, then email lookup is three |
| Price per call | The catalog, shown before each run | What does each endpoint charge? |
| Retries and misses | Your tolerance | How many empty results will you re-run? |

The result is a number you can know before the run starts. Add a margin for retries, since some lookups return nothing and some need a second provider.

[What 1,000 lookups cost](https://looot.ai/blog/what-1000-lookups-cost) walks through this arithmetic with live prices, and [SERP API cost, reading the live catalog](https://looot.ai/blog/serp-api-cost-reading-the-live-catalog) shows how one capability can vary in price by provider.

## Why is the data meter easier to cap?

Because each call has a visible price and a count. You can set the rule in words: no more than this many calls, stop if the balance falls below this. The agent can check both.

The model meter is harder, because token use depends on how long the agent thinks and how much text it reads. You can limit it, but you cannot read it before the run in the same way.

Put differently, a data call is a line item and a token count is a rate. Line items are easier to budget.

## What does looot show before a call?

looot is one connection that gives an agent access to data and tool APIs, with pay per call pricing. Before a run, the agent can call inspect on an endpoint and read its input fields and its price. After a run, the runs tool shows what ran and what each run cost, and the balance tool shows what is left.

If the balance cannot cover a call, the call is refused with a link to top up. That gives you a hard stop at the data meter, not a bill you find later.

The table shows today's prices for three calls an agent often makes. The numbers come from the live catalog and change as providers change.

Price per call for typical agent steps: live prices per call are read from the catalog when the page loads, see https://looot.ai/blog/ai-agent-cost-data-calls.

- Google organic results
- Page as markdown
- Keyword volume and CPC

## How do I keep an agent from overspending?

Four rules cover most cases.

1. **Price first.** Tell the agent to read the price of an endpoint before running it, and to tell you when it passes a limit.
2. **Cap the calls.** Give a maximum number of calls per task. A typical research task needs a handful.
3. **Test small.** Run one item, check the cost, then run the list. A one-item test shows the real price and the real result.
4. **Choose the cheapest provider that returns what you need.** When several providers serve one job, start with the cheapest and move up only when it returns nothing.

You can see the second and third rules at work in [waterfall enrichment explained](https://looot.ai/blog/waterfall-enrichment-explained), where cheaper providers run first.

## Is a free option cheaper?

Not always. A free tier has caps that can stop an agent mid-run, and the time lost to a stalled job costs something too. [Free AI API options and what free means for an agent](https://looot.ai/blog/free-ai-api-for-agents) covers where free tiers break. For a one-off question, a free tool is the cheapest answer. For a list, a priced call you can read in advance is easier to plan.

## A worked estimate

Suppose you want, for a list of 200 domains, the top Google results for each domain and the best page read as markdown. That is two calls per domain, so 400 calls in all.

Read the price of the cheapest search endpoint and the cheapest page endpoint from the table. Call them S and P. The data cost is 200 times S plus 200 times P. Add a margin of ten percent for retries and for the pages that fail to load, and you have the figure to compare against your balance.

Model cost sits on top of that figure. The agent reads each result and writes a short summary, so the number of tokens grows with the length of the pages it reads. If you ask it to read only the first part of each page, the model cost drops, and the data cost stays the same.

## What do retries cost?

A retry is a new call at the full price. If one in ten calls fails and the agent retries each failure once, you pay for ten percent more calls. If the agent retries until it succeeds, there is no limit, and a broken input can cost a lot. That is why the instruction should say how many retries are allowed.

A cheaper habit is to separate transient failures from permanent ones. A timeout is worth one retry. An empty result for a valid input is usually final, so move to the next provider or mark the item as not found. An error that names a bad input will fail again, so fix the input first.

## Try it: a prompt for your agent

This prompt produces a cost estimate for a job you describe, without running anything.

You get a number before any money moves.

## What it cannot do

- It does not show or cap model token cost. That meter belongs to your model vendor.
- It does not predict misses. An empty result still happened at some price, so keep a margin.
- It does not guarantee a price will stay the same. Read the price at the time of the run.
- It does not stop a job on a timer. Anything that needs scheduled runs is not available yet.

Last checked 2026-10-03 against the live catalog and the looot skill file. Author: the looot team.

## Questions

### How much does an AI agent cost?

It depends on the job. Cost is model tokens plus tool calls, and hosting if you run it yourself. Estimate tool calls by multiplying items, calls per item and price per call.

### What costs more, the model or the tools?

It varies by job. A long reasoning task is mostly model cost. A list job with many lookups is mostly tool cost. Estimate both before you start.

### How do I estimate AI agent cost before running?

Count the items and the calls per item, then multiply by the price of each call. Add a margin for retries. With looot the agent can read every endpoint's price without running it.

### Can I set a spending limit for an agent?

Yes. Tell the agent a maximum number of calls and a balance threshold, and have it read prices first. With looot a call that the balance cannot cover is refused.

### Are AI agent data calls priced per call?

With looot, yes. Each endpoint has a price per call, shown before you run it, and there is no subscription for the calls themselves.

### Why did my agent cost more than I expected?

Usually more calls than planned, retries after misses or a loop. Cap the number of calls and read the run history afterwards to see which steps cost the most.

## For agents

### Estimate a job before running it

```text
Use looot to estimate the data cost of this job: find the top five Google results for 40 keywords and read the best page for each as markdown.
1. Set up https://looot.ai/skill.md.
2. Use discover and inspect to find the cheapest endpoint for each of the two steps. Tell me each price.
3. Work out the total: 40 searches plus 40 page reads. Show the arithmetic.
4. Add 10 percent for retries and misses and give me the final figure.
5. Tell me my balance and whether it covers the job. Do not run anything.
```
