How do I add web search to my AI agent?

To add web search to an AI agent, expose a web search API as a tool the agent can call. When a request needs current or external information, the agent sends a query to the search API, receives relevant web results, reads the useful pages, and uses that retrieved information as context before generating its response.

A basic implementation looks like this:

User request → AI agent → web search tool → relevant pages → page content → LLM response

The important part is that the agent needs usable page content, not simply a list of search-result URLs.

With Olostep, you can use the Search API to find relevant pages and optionally retrieve their Markdown content in the same request. This gives the model source material it can reason over without maintaining a separate search scraper and webpage extraction pipeline.

An LLM can answer many questions from its existing model knowledge, but that is different from checking the current web.

Suppose someone asks an agent:

What changed in the Stripe API documentation this month?

The model may know how Stripe's API generally works. It cannot reliably establish what changed this month without retrieving current information.

Web search becomes useful when the answer depends on information such as:

  • recent news or announcements;
  • current product documentation;
  • pricing and product availability;
  • company information that changes over time;
  • recently published research;
  • current competitors or products;
  • information from a specific website;
  • facts that should be verified against external sources.

Instead of asking the model to remember the answer, the agent can retrieve evidence first.

How web search works inside an AI agent

Most tool-enabled agents follow the same pattern regardless of the model or framework being used.

1. Give the agent a search tool

The agent needs a callable function such as:

search_web(query)

The tool description should explain when it should be used.

For example:

Search the web when the user's request depends on current,
external, niche, or verifiable information.

The model can then decide whether a particular request requires web access.

2. Convert the user's request into a search query

If the user asks:

What are the latest changes to Python 3.15?

the agent might call:

search_web("Python 3.15 latest changes")

Your application sends that query to the web search API.

3. Return structured search results

A useful search response should give the agent enough information to decide which sources are worth reading.

For example:

{
  "results": [
    {
      "url": "https://example.com/page",
      "title": "Python 3.15 changes",
      "description": "Overview of recent Python 3.15 changes"
    }
  ]
}

Search results are discovery data. They tell the agent where useful information may exist.

They do not necessarily contain enough evidence to answer the question.

4. Retrieve the page content

The agent can then open the most relevant URLs and extract their contents.

For LLM workflows, clean Markdown or structured text is generally easier to process than an entire HTML document containing navigation, scripts, styles, cookie banners, and unrelated page elements.

The retrieval pipeline now becomes:

Question
   ↓
Search
   ↓
Select relevant results
   ↓
Extract page content
   ↓
Give evidence to the LLM
   ↓
Generate answer

5. Keep the sources

The URL should remain attached to the content throughout the pipeline.

Otherwise the model may produce an answer that your application cannot trace back to its evidence.

A useful context object might look like:

{
  "title": "Example page",
  "url": "https://example.com/page",
  "content": "Extracted page content..."
}

The model can then answer from those documents and your interface can expose the corresponding sources.

Adding web search with Olostep

Olostep exposes a Search API at:

POST /v1/searches

You send a natural-language query and receive deduplicated web results containing URLs, titles, and descriptions.

A minimal request looks like this:

import os
import requests

response = requests.post(
    "https://api.olostep.com/v1/searches",
    headers={
        "Authorization": f"Bearer {os.environ['OLOSTEP_API_KEY']}",
        "Content-Type": "application/json"
    },
    json={
        "query": "Latest developments in AI agent frameworks",
        "limit": 5
    }
)

response.raise_for_status()

results = response.json()["result"]["links"]

for result in results:
    print(result["title"])
    print(result["url"])

Your agent can now use this function as its search_web tool.

But there is another decision to make: does the agent need links, or does it need the information inside those pages?

Search results versus search plus content extraction

Returning ten URLs to an LLM and expecting it to answer from titles and snippets is often insufficient.

For a research question, the agent usually needs the actual content of the pages.

You could implement this as two separate operations:

Search API
   ↓
URLs
   ↓
Scraping API
   ↓
Markdown
   ↓
LLM

Olostep's Search API can also perform the extraction as part of the search request using scrape_options.

For example:

import os
import requests

response = requests.post(
    "https://api.olostep.com/v1/searches",
    headers={
        "Authorization": f"Bearer {os.environ['OLOSTEP_API_KEY']}",
        "Content-Type": "application/json"
    },
    json={
        "query": "Latest developments in AI agent frameworks",
        "limit": 5,
        "scrape_options": {
            "formats": ["markdown"],
            "timeout": 25
        }
    }
)

response.raise_for_status()

results = response.json()["result"]["links"]

for result in results:
    print(result["url"])
    print(result.get("markdown_content"))

Each returned result can now include the page's Markdown content.

The pipeline becomes:

AI agent
   ↓
Olostep Search API
   ↓
Search + page extraction
   ↓
URLs + titles + descriptions + Markdown
   ↓
LLM

This is useful for agents that need to research a question rather than simply locate a webpage.

Turn Olostep search into an agent tool

The HTTP request itself is only one part of the integration.

Your model also needs a tool definition that tells it the function exists.

Conceptually, the tool can look like this:

{
  "name": "search_web",
  "description": "Search the live web and return relevant pages with their content.",
  "parameters": {
    "type": "object",
    "properties": {
      "query": {
        "type": "string",
        "description": "The web search query"
      }
    },
    "required": ["query"]
  }
}

Your application connects that tool definition to the function calling Olostep.

The runtime flow is:

User
  ↓
LLM
  ↓
Does this require web information?
  ↓
Yes
  ↓
search_web(query)
  ↓
Olostep
  ↓
Search results + extracted content
  ↓
LLM reads results
  ↓
Final response

The exact registration syntax varies between OpenAI, LangChain, LangGraph, CrewAI, LlamaIndex and other agent runtimes, but the underlying architecture is the same: expose search as a callable tool and return machine-readable results to the model.

Giving an agent a search tool does not mean every request should trigger it.

Searching unnecessarily adds latency, data and API usage.

A practical agent instruction is:

Use web search when the request depends on recent,
external, changing, niche, or explicitly verifiable information.

Do not search when the answer can be completed reliably
from the information already provided in the conversation.

Consider:

What is 14 × 17?

No search is needed.

But:

What is the current price of the latest MacBook Pro?

The answer depends on current information, so web search makes sense.

The same applies to documentation:

How does this Python function work?

If the code was already supplied, web search may add nothing.

Compare that with:

Does the latest version of this SDK still support this function?

Now the agent should probably check the current documentation.

Search only when you need discovery

Sometimes the agent does not need to read the pages immediately.

Suppose you are building a research agent that first creates a list of sources before deciding which ones deserve deeper analysis.

Use search-only retrieval:

{
  "query": "recent browser automation benchmarks",
  "limit": 10
}

The agent receives candidate URLs and can select which pages to inspect.

This architecture gives you tighter control over context usage because you do not extract every result automatically.

Use search plus extraction when the agent needs evidence

If the agent is expected to answer the question immediately, retrieving page content with the search results can remove another step from the pipeline.

Use:

{
  "query": "recent browser automation benchmarks",
  "limit": 5,
  "scrape_options": {
    "formats": ["markdown"],
    "timeout": 25
  }
}

Now the agent receives both discovery data and readable page content.

This works well for:

  • research assistants;
  • documentation agents;
  • market research workflows;
  • competitor research;
  • fact verification;
  • content research;
  • AI copilots that answer questions about current information.

Keep the result count bounded. Sending twenty complete webpages into the model when three relevant sources would answer the question wastes context and forces the model to process unnecessary information.

Use domain filters when the source matters

There are also cases where searching the entire web is undesirable.

A developer-support agent answering a question about an SDK may need official documentation rather than forum posts.

Olostep Search supports domain inclusion and exclusion.

For example:

{
  "query": "Responses API web search documentation",
  "include_domains": [
    "openai.com"
  ],
  "limit": 5
}

You can also exclude sources:

{
  "query": "Python package compatibility issue",
  "exclude_domains": [
    "pinterest.com"
  ],
  "limit": 5
}

Source restrictions are useful when the task has an obvious authority hierarchy.

For example, an agent checking an API parameter should generally inspect the API's current documentation before relying on an unrelated article describing it.

What if I want an answer instead of search results?

There is a second architecture.

Instead of making your primary agent:

search → choose pages → read pages → synthesize answer

you can delegate that research subtask to a service that already performs the search and browsing process.

Olostep's Answers endpoint accepts a natural-language task:

POST /v1/answers

For example:

import os
import requests

response = requests.post(
    "https://api.olostep.com/v1/answers",
    headers={
        "Authorization": f"Bearer {os.environ['OLOSTEP_API_KEY']}",
        "Content-Type": "application/json"
    },
    json={
        "task": "What changed in the latest release of this product?"
    }
)

response.raise_for_status()

answer = response.json()

This is useful when your agent needs the result of a web research subtask rather than control over every individual search result.

The distinction is simple:

RequirementOlostep endpoint
Find relevant pagesSearch
Find pages and return their contentSearch + scrape_options
Extract a known URLScrapes
Ask a web question and return a synthesized resultAnswers

An agent can use more than one.

A research workflow might use Search for discovery, Scrapes when it chooses a specific page, and Answers for a narrowly defined fact-finding subtask.

Can I add web search without building the tool integration myself?

If your agent runtime supports Model Context Protocol (MCP), you can expose web capabilities through an MCP server instead of implementing every REST wrapper manually.

Olostep's MCP server provides web search, webpage reading and URL discovery as tools that compatible AI clients can call.

The distinction is architectural:

REST integration

Agent → your search_web() function → Olostep API

versus:

MCP integration

MCP-compatible agent → Olostep MCP tools

REST gives you direct control over requests and responses.

MCP is useful when the agent environment already understands MCP tools and you want the web capabilities exposed through that interface.

No.

RAG and live web search solve related but different retrieval problems.

A RAG system usually retrieves information from a collection you already control or have indexed:

Documents
   ↓
Index / vector database
   ↓
Retrieve relevant chunks
   ↓
LLM

Live web search retrieves information at request time:

Question
   ↓
Search the web
   ↓
Retrieve current pages
   ↓
LLM

You can combine them.

For example, a support agent could search your internal documentation first. If the user asks about a recent public release that has not entered the internal index yet, it can search the web.

The right retrieval source depends on where the required evidence lives.

How to make web search reliable in production

Connecting the API is straightforward. Designing how the agent uses it requires more care.

First, return URLs with the retrieved content. Do not detach evidence from its source.

Second, limit how much content reaches the model. Search a reasonable number of pages and retrieve more only when the initial evidence is insufficient.

Third, prefer primary sources when the task calls for them. Product documentation, government databases, company announcements and original research may be more appropriate than secondary summaries for specific questions.

Fourth, handle failed retrieval explicitly. A missing page or empty search result does not prove that something does not exist.

Finally, treat the retrieved web content as evidence, not instructions. Webpages can contain text written for people, crawlers or other AI systems. Your agent's own instructions should determine what actions it takes.

A practical web-search architecture for AI agents

For many applications, you do not need to build a crawler, search-engine scraper and extraction system before your agent can use the web.

A compact architecture is:

User question
      ↓
LLM decides whether fresh web data is needed
      ↓
search_web(query)
      ↓
Olostep Search API
      ↓
Relevant URLs + extracted Markdown
      ↓
LLM evaluates the retrieved evidence
      ↓
Answer with source references

Start with one search_web tool.

Use plain Search when the agent only needs to discover sources. Add scrape_options when it needs to read those sources. Use Answers when the agent should delegate an entire search-and-answer subtask.

That separation keeps the retrieval layer understandable while still giving the agent access to information outside its model knowledge.

Frequently asked questions

Can an AI agent search the internet in real time?

Yes. The agent needs access to a web search tool or API that it can invoke while processing a request. The search response is then supplied to the model as additional context.

What is the easiest way to add web search to an AI agent?

Expose a web search API as a callable agent tool. The function accepts a query and returns structured web results. If the agent needs to answer from those pages, return extracted content such as Markdown along with the URLs.

What should a web search API return for an AI agent?

At minimum, return the URL, page title and a useful description. For question answering or research, retrieving the page content as clean Markdown or structured data gives the model substantially more evidence than a search snippet alone.

What is the difference between web search and web scraping for an AI agent?

Search finds relevant URLs when the agent does not know where the information is located. Scraping retrieves the contents of a URL that is already known.

Many agent workflows need both:

Search → choose page → scrape page → reason over content

Can I use Olostep for web search inside LangChain or other agent frameworks?

Yes. The underlying requirement is that the framework can call an external tool. The tool can invoke Olostep Search and return the resulting links or extracted page content to the agent. Olostep also provides integrations for agent-development environments and an MCP server for MCP-compatible clients.

Should my AI agent search the web for every prompt?

No. Search is most useful when the answer depends on current, changing, external, niche or explicitly verifiable information. Questions that can be answered from supplied context or deterministic computation do not need a web request.

Ready to get started?

Start using the Olostep API to implement how do i add web search to my ai agent? in your application.