AI Agents
Arslan
ArslanOct 5, 2026

Learn how agentic retrieval works, how it differs from classic RAG, when to use live web search, what tools it needs, and how to control cost and latency.

What Is Agentic Retrieval? How AI Agents Search, Plan, and Verify

Ask a classic RAG system a two-part question, and it runs one search and hopes the right passages come back. Ask an agent with agentic retrieval, and it splits the question, picks a source for each part, checks what it found, and searches again if something is missing.

This guide explains how that loop works, which tools it needs when the answer lives on the live web, what it costs, and when a simpler pipeline is the better choice.

What Agentic Retrieval Means

Agentic retrieval is a retrieval pattern where an LLM agent controls the search process. The agent decides what to look for, which source or tool to use, whether the results are good enough, and when to search again, instead of running one fixed query and passing the top results to the model.

One agentic RAG survey paper describes the idea this way: "Agentic Retrieval-Augmented Generation (Agentic RAG) transcends these limitations by embedding autonomous AI agents into the RAG pipeline." Many writers use "agentic retrieval" and "agentic RAG" for the same pattern.

The term shows up in two ways, which causes confusion:

  • As a pattern: any system where an agent plans, runs, and checks its own retrieval steps. This is the meaning used in this guide.
  • As a product feature: Azure AI Search ships a feature called "agentic retrieval," with its own API and limits.

Search is one step inside the loop. If you want that step on its own, see agentic search explained.

How Agentic Retrieval Differs From Classic RAG

Classic RAG retrieves once, then generates. It turns the user's question into an embedding, pulls the top matching chunks from a vector store, and hands them to the model in a single pass.

Agentic retrieval adds decisions between those steps. That can make it better at complex questions, but slower and more expensive. Microsoft's classic RAG guidance puts the trade-off plainly: "This approach is simpler with fewer components, and faster because there's no LLM involvement in query planning."

Classic RAGAgentic retrieval
Query handlingOne query, used as writtenRewritten or split into subqueries
Retrieval callsOneSeveral, decided at run time
Sources reachedUsually one indexVector store, database, live web, APIs
Data freshnessAs fresh as the last indexing jobCan fetch current pages at query time
Self-checkNoneGrades results and retries
Latency and costLowerHigher, grows with each step
Best forSimple lookups over stable documentsMulti-part, comparative, or time-sensitive questions

How the Agentic Retrieval Loop Works

The loop has five steps, and the agent can repeat steps three to five several times. Take one question as an example: "How do the current free tiers of three vector databases compare?"

Plan the Query

The agent first breaks the question into smaller subqueries. Here, that means one search per vector database, plus a check for the date each pricing page was last updated.

Azure AI Search does this with an LLM that plans subqueries and runs them in parallel; that planning feature is still in preview. The Azure agentic retrieval docs are clear about the cost: "Agentic retrieval adds latency compared to a single-query pipeline, but it handles query complexity that a single query can't."

Route to the Right Source

Next, the agent picks where each subquery should go. A vector store works for private, stable documents, and a SQL database works for structured records.

Pricing pages change without notice, so the example question belongs on the live web. Routing is the step where the agent decides between its own index and a web search.

Retrieve and Read

The agent calls the chosen tools and reads what comes back. For the example, it searches for each vendor's pricing page, then reads those pages.

Each piece of evidence should keep its source URL and the time it was fetched. Without that metadata, the agent cannot cite its answer or tell whether a fact is out of date.

Grade the Evidence

Before answering, the agent checks whether the results are relevant and complete. If one vendor's page didn't load or only covers the paid plan, that subquery fails the check.

Corrective RAG is a named version of this step, which the agentic RAG survey describes. It scores retrieved passages and triggers a new search when they fall short.

Decide to Answer or Search Again

Finally, the agent either writes the answer or loops back with a better query. The stopping rule should depend on evidence coverage and a fixed budget, not on how confident the model sounds.

Multi-step search pays off most on questions that need several hops. In the Search-R1 benchmark results, Qwen2.5-7B with multi-turn search averaged 0.384 Exact Match across seven QA datasets vs 0.304 for one-round RAG. That model was trained with reinforcement learning to search a 2018 Wikipedia corpus, so the gain reflects a trained search policy on static data, not the live web.

Why Agentic Retrieval Needs the Live Web

An agent that loops over a stale index still gets stale answers. Planning and grading can improve which chunks it picks, but they can't add facts that were never indexed.

Many questions agents get depend on data that changes after indexing: prices, documentation, changelogs, job posts, news, and competitor pages. For those, the agent needs a way to search the web and read current pages at query time.

Agentic retrieval is only as good as the sources the agent can reach. If those sources stop at the vector store, so do the answers.

The Tool Kit for Agentic Retrieval Over the Web

A web-facing agent needs a small set of tools, each with a clear job. The agent calls them through function calling, or through a standard interface such as the Model Context Protocol (MCP), which the Olostep MCP server supports.

ToolWhen the agent calls itWhat it should return
SearchIt doesn't know which pages hold the answerRanked URLs with titles and snippets, optionally page content
Scrape a URLIt already knows the pageClean Markdown or JSON for that one page
Crawl a siteIt needs many pages from one domain, such as docsContent for each page, with URLs
Map a siteIt needs to see a site's pages before choosingA list of URLs
Extract structured dataIt needs specific fields, such as plan name and priceJSON that matches a schema

A web search API for agents handles discovery. A scrape endpoint for single pages handles known URLs, and a crawler can crawl an entire site when the answer is spread across a documentation set.

Return Clean Markdown or JSON, Not Raw HTML

The format of retrieved content decides how much of the context window goes to useful text. Raw HTML carries navigation, scripts, and styling that cost tokens and distract the model.

Markdown works best when the agent needs to read and reason over a page. JSON works best when it needs exact fields to compare. This breakdown of LLM-ready web data formats covers the trade-offs in more detail.

A Minimal Search-Then-Scrape Loop

The sketch below shows the loop with two guardrails: a cap on tool calls and a forced final answer. It uses Olostep's search endpoint, which can return each result's page content as Markdown in the same call.

python
import requests

API = "https://api.olostep.com/v1"
HEADERS = {"Authorization": "Bearer <OLOSTEP_API_KEY>"}
MAX_TOOL_CALLS = 8

def search(query):
    r = requests.post(f"{API}/searches", headers=HEADERS, json={
        "query": query,
        "limit": 5,
        "scrape_options": {"formats": ["markdown"]},
    })
    return r.json()["result"]["links"]  # url, title, description, markdown_content

def agentic_retrieve(question, plan_subqueries, is_sufficient, answer):
    evidence, seen_urls, calls = [], set(), 0
    queue = plan_subqueries(question)            # LLM step: split the question
    while queue and calls < MAX_TOOL_CALLS:
        query = queue.pop(0)
        calls += 1
        for link in search(query):
            if link["url"] in seen_urls or not link.get("markdown_content"):
                continue                         # skip duplicates and failed pages
            seen_urls.add(link["url"])
            evidence.append({"url": link["url"], "text": link["markdown_content"]})
        gaps = is_sufficient(question, evidence) # LLM step: grade coverage
        if not gaps:
            break
        queue.extend(gaps)                       # search again for what's missing
    return answer(question, evidence)            # final answer, no more tools

The three LLM steps are left as functions so you can plug in the model of your choice. If you'd rather not run the loop yourself, the Answers endpoint takes a task, searches and reads the web, and returns an answer with its sources, optionally as structured JSON that matches a schema you provide.

What Agentic Retrieval Costs

Every extra step in the loop adds model calls, tokens, and waiting time. Anthropic measured this in its own research system, and the Anthropic multi-agent research post reports that "agents typically use about 4× more tokens than chat interactions, and multi-agent systems use about 15× more tokens than chats."

The same post reports that its multi-agent setup beat a single agent by 90.2% on an internal research eval. That gain came with the 15× token cost, so judge quality and spend together.

A rough per-query estimate is simple: number of LLM calls × tokens per call × price per token, plus the cost of each search or scrape. These failure modes are common sources of waste:

Failure modeWhat you seeFix
Runaway loopDozens of tool calls for one questionCap tool calls, then force an answer with no tools
Duplicate fetchesThe same URL read several timesTrack seen URLs and cache evidence
Context bloatLong prompts, slower and worse answersReturn Markdown or JSON, trim to relevant sections
Stale indexConfident answers with old factsRoute time-sensitive subqueries to live web search
Failed or empty pagesMissing evidence for one subqueryRetry once, then mark the gap in the answer

When to Skip Agentic Retrieval

Skip agentic retrieval when one search can answer the question. Simple lookups over stable documents get faster, cheaper, and more predictable results from classic RAG.

The agentic RAG survey makes the same point: "Agentic RAG should not be viewed as a universal replacement for traditional RAG." Gartner flags a wider risk. The Gartner agentic AI prediction forecasts that "Over 40% of agentic AI projects will be canceled by the end of 2027, due to escalating costs, unclear business value or inadequate risk controls."

A practical approach is to route each query:

  • Use classic RAG for single-fact questions over documents that rarely change.
  • Use agentic retrieval for multi-part, comparative, or multi-hop questions.
  • Use agentic retrieval with live web tools when the answer depends on data that changes, such as prices, releases, or news.
  • Use neither when a fixed scraper on known pages will do; a deterministic job is usually cheaper than an agent.

How to Evaluate an Agentic Retrieval System

Measure the path the agent took as well as the final answer. Two systems can give the same answer while one used three tool calls and the other used thirty.

Track these numbers on a fixed test set:

  • Answer accuracy: is the final answer correct?
  • Citation correctness: does each cited URL support the claim attached to it?
  • Tool calls per answer: how many searches and scrapes did it take?
  • Cost per correct answer: total spend divided by correct answers.
  • Freshness: include questions whose answers changed recently, so the test catches a stale index.

FAQ

Is Agentic Retrieval the Same as Agentic RAG?

Usually. Most articles use the two terms for the same pattern: an agent that plans, runs, and checks its own retrieval before generating an answer.

Is RAG Agentic AI?

Classic RAG is not agentic, because it runs one fixed retrieval step. It becomes agentic when an LLM decides what to retrieve, from where, and whether to retrieve again.

What Are RAG Agents?

RAG agents are LLM agents that use retrieval as a tool. They choose when to search a vector store, database, or the web, then use the results to answer.

Is Agentic Retrieval a Product or a Pattern?

It is a pattern first. Azure AI Search also offers a feature with that name, with its own API and limits.

Can Agentic Retrieval Search the Live Web?

Yes, if you give the agent web tools such as search, scrape, and crawl. Without them, the agent can only retrieve what is already in its index.

How Many Tool Calls Should an Agent Make?

Set a hard cap based on your latency and cost budget; under ten calls per question is a common starting point. When the cap is hit, force a final answer and note any gaps.

Should I Use One Agent or Several?

Start with one agent. Add more only when a task splits into independent parts, because each extra agent adds tokens and coordination time.

How Do You Keep Sources Cited Across Steps?

Store the source URL and fetch time with every piece of evidence. Pass that metadata into the final prompt so the model can cite each claim.

About the Author

Arslan Ali

Co-Founder, Olostep · San Francisco, CA

Arslan is the co-founder of Olostep, a web data infrastructure platform that helps developers and teams access, extract, and structure web data at scale. He works closely on the product and technology behind Olostep, with a focus on building reliable infrastructure for web scraping, search APIs, and structured web data.

Read more