Tutorial
Arslan
ArslanOct 6, 2026

Learn what data freshness means, how to measure data age and staleness, set freshness SLAs, and keep web data and RAG systems up to date.

What Is Data Freshness? Definition, Metrics, and Best Practices

A dashboard can load quickly, a pipeline can report success, and an AI agent can answer confidently, and the data underneath can still be hours or days out of date. Data freshness is how you tell the difference. This guide defines the term, separates it from look-alike terms, shows how to measure it, and explains why it matters more once AI systems start acting on external web data.

What Is Data Freshness?

Data freshness is how current your data is at the moment you use it. You measure it as the time between when something changed at the source and when your system reads it. If a product price changed at 9:00 and your agent reads it at 9:45, that data is 45 minutes old.

Fresh does not mean "zero seconds old." Data is fresh when its age is small enough for the job it is doing. A 45-minute-old price is fine for a weekly market report. It is not fine for an agent about to place an order.

Other Names for Data Freshness

You will see a few other terms used for the same idea:

  • Data currency: how well data reflects the current state of the world.
  • Data recency: how recently the data was collected or updated.
  • Up-to-dateness: a plain-language way to describe the same thing.
  • Data age: the measured number, such as "45 minutes" or "3 days."

Data quality frameworks such as the UK government's group freshness under a broader quality dimension called timeliness. The next section explains how these terms differ.

Data Freshness vs. Latency, Timeliness, and Recency

These four terms often get mixed up. The simplest way to separate them is to ask what each one describes: a state, a process, a requirement, or a timestamp.

TermWhat It DescribesQuestion It AnswersExample
Data freshnessA state: the age of data when it is used"How old is this value right now?"The stock level your agent sees is 20 minutes old
Data latencyA process: the delay added at each step of a pipeline"How long does data take to move from source to use?"Extraction takes 5 minutes and indexing takes 10
Data timelinessA requirement: data is ready when the use case needs it"Is it available by the deadline?"The pricing report needs data by 8:00 every morning
Data recencyA timestamp: when data was last collected"When did we last fetch this?"The page was scraped at 06:12 UTC

The UK government frames timeliness around the intended use. The UK Government Data Quality Framework states: "Data is timely if the time lag between collection and availability is appropriate for the intended use."

Low latency does not guarantee fresh data. Some vendors use "latency" to mean query speed. A query can return in 50 milliseconds and still serve a value that was collected last week. Speed measures the pipe, and freshness measures what flows through it.

What Does It Mean for Data to Be Stale?

Stale data is data whose age is greater than your use case can tolerate. The danger is that stale data looks normal. The row exists, the field is filled in, and nothing in the value itself tells you it is out of date.

Data goes stale for a few common reasons:

  • Failed or skipped jobs: a scheduled refresh does not run, so old values stay in place.
  • Jobs that succeed but fetch nothing new: the run reports success, but the source returned old or empty results.
  • Caches serving old copies: a cached response is returned long after the source changed.
  • Source changes that break extraction: a website changes its layout, the scraper misses the new field, and the old value is never replaced.
  • Slow upstream sources: a partner feed or public dataset updates later than you expected.

There is also data decay. Even with a perfect pipeline, facts keep changing after you collect them. Prices change, people switch jobs, and policies get revised. A record that was correct when you collected it becomes wrong with time.

How to Measure Data Freshness

You cannot measure freshness without timestamps. Most systems track up to three:

  • Event time: when the change happened at the source.
  • Fetch or ingestion time: when your system collected the data.
  • Availability time: when the data became ready to query or use.

Event time is the most accurate reference, but it is often unknown. In that case, teams use fetch time as a stand-in and accept that the true age may be higher.

Core Freshness Metrics

These formulas cover most monitoring needs:

text
Data age            = now − source change time   (use fetch time if change time is unknown)
Freshness lag       = availability time − event time
Average / p95 age   = mean or 95th-percentile data age across all records
Staleness ratio     = records older than the threshold ÷ total records
SLA attainment      = time (or records) within the freshness target ÷ total
Expected age with a fixed refresh schedule ≈ (refresh interval ÷ 2) + processing time
Worst-case age      ≈ refresh interval + processing time

Here is a worked example. Say you scrape a set of competitor pricing pages every 6 hours, and processing takes 15 minutes. The average record will be about 3 hours and 15 minutes old, and the oldest will be about 6 hours and 15 minutes old. If your freshness target is 4 hours, a large share of your records will miss it on any given day. Your schedule, not your pipeline, is the problem.

A data freshness SLA (service level agreement) is simply a target written down, such as "95% of records under 4 hours old." The p95 age and staleness ratio tell you whether you are meeting it.

The Two Clocks of Web Data

External web data has two clocks. The first clock is when the page changed. The second is when you fetched it.

A page fetched one minute ago can show a price that was set three days ago. That is fine, because your copy matches the source. The real problem is the opposite case: your copy is three days old, and the page changed this morning.

To catch that gap, track both the fetch time and the last detected change for every page. Useful signals include HTTP Last-Modified headers, lastmod dates in sitemaps, and content hashes that show whether the page text actually changed. A workflow for web scraping change tracking formalizes this: take a baseline, scrape on a schedule, compare the results, and alert on differences.

How Fresh Does Data Need to Be?

Fresher is not always better. Every refresh costs compute and money, and fetching live data takes longer than reading a cached copy. The right target is the one that fits the intended use, not the lowest number you can reach.

Decide using three inputs:

  • How fast the source changes: stock levels can change by the minute, while API reference docs may change monthly.
  • The cost of acting on a stale value: a wrong price in an automated order costs more than a stale fact in a weekly report.
  • The cost of a refresh: fetching 10 pages every minute is cheap, while fetching 1 million pages every minute is not.
Freshness TierTypical Age TargetWeb and AI Examples
Seconds to minutesUnder 5 minutesAn agent checking stock, price, or availability before it acts
Hourly1 hourPricing intelligence, news and policy monitoring
Daily24 hoursLead and company enrichment, competitor page tracking
Weekly or longer7+ daysRAG over stable reference content, such as documentation archives

Different fields in the same dataset can need different tiers. A company's headcount can be a week old, while its pricing page may need hourly checks.

Why Data Freshness Matters for AI Agents and RAG

AI systems make stale data more dangerous because they state outdated facts in the same confident tone as current ones. In an AI application, staleness has two sources: what the model learned in training, and what you retrieve for it at answer time.

Model Knowledge Is Stale by Design

A large language model (LLM) learns from training data collected up to a cutoff date. After that date, it knows nothing new unless you give it outside information. The LLMLagBench study puts it directly: "This creates a strict knowledge boundary beyond which models cannot provide accurate information without querying external sources."

The authors also note that provider-declared cutoffs and what a model actually knows can differ, so the stated date is only a rough guide. For any question that depends on current facts, an agent needs live, external data. The model's memory alone is not enough.

Stale Retrieval Can Be Worse Than No Retrieval

Retrieval-augmented generation (RAG) adds documents to a model's prompt so it can answer from current sources. But RAG only helps if those documents are still valid. Research on stale-document poisoning found that "outdated retrieval flips 30% of Llama and 37% of Qwen answers even without instructions to trust the document." In those cases, the model would have answered correctly without the outdated document.

Adding new documents does not fix this on its own. When a fact changes, the index often holds both the old and new versions, and they look almost identical to the retriever. A 2026 retrieval memory benchmark reported that "when required to answer, RAG serves superseded values 15-40% of the time." Peer-reviewed work on the HoH benchmark reached a similar conclusion: "outdated information significantly degrades RAG performance."

The practical takeaways for RAG freshness:

  • Retire superseded records: remove or mark old versions when a source changes, instead of only adding new chunks.
  • Store reliable timestamps: keep fetch time and source change time in each chunk's metadata.
  • Rank with recency carefully: in the stale-document poisoning study, a recency-aware re-ranker reduced errors only when document dates were accurate.

How to Keep External Web Data Fresh

Keeping web data fresh comes down to matching the refresh method to the freshness tier. Three patterns cover most cases.

Fetch Live When the Answer Depends on Now

For decisions on fast-changing facts, fetch the page at decision time. A live single-URL scrape returns the current page as clean Markdown, HTML, text, or structured JSON, so the agent works from the source as it looks now.

Live fetching has a cost: it takes seconds, while a cached result returns in milliseconds. The guide to cached versus live scraping compares the two, with 2–10 seconds for a live scrape and 50–200 milliseconds for a cached one. Olostep's maxAge parameter sets the maximum acceptable age of cached data for each request. According to Olostep's caching behavior docs, the default cache window is two days, and caching can speed up scraping by up to 5x for content that does not need real-time updates.

This lets each request choose its own freshness. A price check can require data no older than a few minutes, while a documentation lookup can accept a day-old copy.

Refresh on Change, Not Only on a Timer

Fixed schedules have two failure modes. They waste fetches on pages that did not change, and they miss changes that happen between runs. Change-driven refresh reduces both.

Use scheduled website monitors to watch high-value pages and alert you when something changes. Then re-collect only the changed pages using a batch refresh of URL lists, which processes many URLs as a single job. Refresh spending then follows the actual change rate of your sources.

Timestamp, Retire, and Monitor

Freshness you cannot see is freshness you cannot manage. Store the source URL, fetch time, and last detected change on every record. When a source changes, retire the old version in your index so it cannot be retrieved again.

Then monitor the data itself, not just the jobs. A job can succeed while the data it delivers is stale. Alert on staleness ratio and p95 age against your target. AI agents can run this loop continuously, as described in Olostep's guide to agent-led dataset maintenance: detect changes, refresh what changed, validate the result, and escalate unclear cases to a person.

Frequently Asked Questions

What Does "Refreshing Data" Mean?

Refreshing data means re-collecting or recomputing it from the source so stored values match the current state. A refresh can run on a schedule, on demand, or when a change is detected.

Is Data Freshness Part of Data Quality?

Yes. Freshness is usually treated as part of the timeliness dimension of data quality, alongside accuracy, completeness, consistency, validity, and uniqueness.

What Causes Poor Data Freshness?

The most common causes are failed or skipped jobs, jobs that succeed without fetching new data, caches that serve old copies, and source changes that break extraction. Refresh schedules that are slower than the source's change rate also cause it.

How Often Should You Re-Scrape a Website?

Re-scrape a page at least as often as it meaningfully changes and as often as your use case requires, whichever is stricter. For large page sets, change monitoring is usually cheaper than re-scraping everything on a fixed timer.

How Do You Keep a RAG Knowledge Base Fresh?

Re-collect changed sources, retire superseded chunks instead of only adding new ones, and store fetch and change timestamps in metadata. Then track the age of retrieved documents so you can see when the index falls behind.

About the Author

Arslan Ali

Co-Founder, Olostep · San Francisco, CA

Arslan is the co-founder of Olostep, a web data infrastructure platform that helps developers and teams access, extract, and structure web data at scale. He works closely on the product and technology behind Olostep, with a focus on building reliable infrastructure for web scraping, search APIs, and structured web data.

Read more