Setup

Claude RAG with live data from an API

Claude RAG usually means a vector store. For facts that change, retrieving from a live API works better. How the two differ and how to wire Claude to live data.

Walid Boulanouar · · 7 min read

On this page

Claude RAG usually means this: you split your documents into chunks, store them in a vector database, and when a question arrives you fetch the closest chunks and hand them to Claude with the question. It is a good design for a fixed pile of text, such as your docs or a contract archive.

It is a weaker design for facts that change or that you never stored. A competitor's current search results, a company's latest headcount or today's version of a web page do not live in your vector store. For those, retrieval means calling a live API at question time. This post explains the difference and shows how to set up live retrieval with Claude.

What is RAG with Claude?

RAG stands for retrieval-augmented generation. Retrieval finds the information, and generation writes the answer from it. Claude does the generation. Where the information comes from is up to you.

The question is where the retrieval step reads from. There are two common sources. A store you filled earlier, which is fast and private but only as fresh as your last update. Or a live source called at question time, which is fresh but costs a call and depends on the provider being up.

Neither wins in general. Use the one that matches how the facts behave.

When should retrieval read from a store?

Use a store when the content is yours, changes slowly and is queried often. Your product docs, support history and internal notes fit. You pay the cost of indexing once, and every question after that is cheap and fast.

A store also gives you control. You decide what goes in, you can remove a document, and the data stays in your systems. For private or regulated material that matters more than freshness.

When should retrieval read from a live API?

Use a live API when the facts belong to the outside world and change. Search results, a company's current details, a page on the web, a social profile, a keyword's current volume. Storing these means copying someone else's changing data and keeping it up to date, which is a job in itself.

Question typeBetter sourceWhy
What does our refund policy say?StoreYour text, changes rarely
Who ranks for this query today?Live APIOutside data, changes
What does this company do?Live APIRead the page when asked
Which of our past tickets mention billing?StoreYour data, queried often
What is this keyword's search volume?Live APIProvider data, one value

A live call costs money each time and adds delay. Cache the answer in your own store when you will ask the same thing again.

How do I give Claude live retrieval?

Claude calls tools. If you connect it to a tool that searches the web and a tool that reads a page, it can retrieve as it answers. The model decides when to search, reads the results, fetches the best page and writes with the text in front of it.

You can do that in two ways. Write your own tool functions and handle the keys, rate limits and failures yourself. Or connect Claude to a server that already offers them. Claude Code and Claude Desktop both connect to MCP servers, and the Claude MCP servers for data work post covers the setup.

With looot, Claude connects once and uses discover to find a search or page-reading endpoint, reads its price with inspect and runs it. The live table shows what those two calls cost today.

For provider-specific setup, the Exa MCP page and the Firecrawl MCP page list their endpoints.

How do I keep live RAG accurate?

Three rules keep answers honest.

Cite the source. Tell Claude to name the URL for every fact it takes from retrieval. If it cannot, the fact should not be in the answer.

Keep the text short. A whole page in markdown can be thousands of words. Ask Claude to pull only the part it needs, or to read the top two pages instead of ten.

Say when retrieval fails. If a page does not load or a search returns nothing, the answer should say so, not fill the gap from memory. Write that into the instruction.

Cost stays predictable if you set a limit on the number of calls per question. Two searches and two page reads is a reasonable ceiling for most questions.

How do I combine a store and a live API?

Most useful setups use both. The store holds what you own, and the live API fills in what changes. A support assistant is a good example. The store answers "what is our refund policy". The live API answers "what is this customer's company doing this quarter", by reading their site when asked.

Tell Claude which source to use for which kind of question. Write it as a short rule in the instruction: policy questions come from the store, questions about outside companies or pages come from a live search or page read. Without the rule the model may search the web for something that sits in your own docs.

When a live answer is worth keeping, write it into your store with its source URL and the date. The next question then costs nothing, and you can see how old the fact is. A fact with no date is a risk, because it will be used long after it stopped being true.

A worked example

Suppose a salesperson asks Claude what a prospect's company sells and who its competitors are. A store of your own notes would be empty for a company you have never met. A live retrieval works like this. Claude searches for the company name, reads the homepage and one product page as markdown, and writes a short summary with two source links. If the pages do not name competitors, it says so, and offers one more search.

That is four calls in total: two searches and two page reads. The cost is the sum of their listed prices, which you can read in the table above before you run anything. Compare that with the cost of building and maintaining a store of every company a salesperson might ask about, and the live call is the cheaper design for questions you cannot predict.

Try it: a prompt for your agent

This prompt runs one live retrieval and answer with sources. Replace the question with one of your own.

Prompt for your agent

Answer with live retrieval

Answer with live retrieval
Use looot to answer a question with live retrieval. 1. Set up https://looot.ai/skill.md. 2. Use discover to find a web search endpoint and a page-as-markdown endpoint. Tell me each price and wait for me to say go. 3. After I say go, search for "what is retrieval augmented generation" and read the two most relevant pages. 4. Write a 150 word answer. After each fact, give the URL it came from. If a page failed to load, say so. 5. Tell me how many calls you made, what they cost and my balance. Use at most four calls.

The agent shows prices first and stops at four calls.

What it cannot do

  • looot does not store or index your documents. It has no vector database or embeddings. Use your own store for private text.
  • It does not make Claude's answer correct. Retrieval gives Claude evidence, and the answer still needs a check.
  • It does not cache results for you. Ask the agent to save an answer to your own notes if you will reuse it.
  • It does not refresh answers on a timer. Keeping an answer current needs scheduled runs, not available yet.

Last checked 2026-10-03 against the live catalog and the looot skill file. Author: Walid.

Questions

Claude RAG is retrieval-augmented generation with Claude as the model. A retrieval step finds relevant information, and Claude writes an answer from it. The retrieval can read from a store you built or from a live API.

ShareXLinkedInEmail
Setup

What is an MCP server? A plain explanation

An MCP server gives an AI agent tools it can call. What a tool call looks like, and how one connection can reach many providers.

Walid Boulanouar · · 5 min read

Comparisons

MCP vs API for agents: when each one fits

MCP and a plain API both let an agent reach outside data. What the agent handles in each, when to pick which, and a prompt to test both on your job.

Walid Boulanouar · · 5 min read

Try it with the agent you already use

Start your workspace, top up, and paste one prompt into your agent.