RAG Corpus Refresh: Reindexing, Re-Embedding, and Freshness Operations

NoraLin 12 2026-10-06 20:42:18 Edit

A RAG system is only as current as its corpus, and corpora rot quietly: documents change, connectors break, thresholds drift, and the system keeps answering — confidently, from the past. Most teams discover staleness through a user complaint about an answer everyone knows is outdated. The alternative is an operated refresh pipeline: triggers that are written down, incremental mechanics that avoid brute force, and tests that catch silence before users do.

Prerequisites: Know Your Two Triggers

Two triggers govern the whole pipeline: content drift — the documented threshold sits around ten to fifteen percent of the corpus changed, before which incremental indexing handles everything and after which re-embedding pays — and embedding-model change, the non-negotiable event where every vector must be recomputed because embeddings from different models are not comparable and a half-migrated index answers from two incompatible geometries at once.

TriggerThresholdWhat it forces
Content drift~10-15% of corpus changedFull re-embed pays; below it, incremental suffices
Embedding-model changeAny switch, any timeMandatory full re-embed — vectors are not cross-compatible
Ordinary change flowContinuousIncremental indexing on arrival; no ceremony

The model-change trigger deserves its non-negotiable status in plain terms: with the same embedding model, re-indexing is unnecessary; the moment the model changes, everything must be re-embedded — because similarity search across two embedding spaces is arithmetic without meaning, and the failure mode is retrieval that looks fine and answers wrong.

Build the Refresh Pipeline: Incremental, Not Brute Force

Build incremental, not brute force: changes flow through an incremental indexing path as they land (new and updated chunks re-embedded on arrival, structured live data queried directly rather than indexed at all), full re-embedding reserved for the two triggers — and the pipeline treats the model-change re-embed as a migration project with dual-index cutover, because the expensive periodic full reindex that brute-forces freshness is the documented anti-pattern the cheaper design exists to replace.

  1. Incremental on change: changed documents flow to re-embedding on arrival — the corpus stays current without scheduled ceremonies.
  2. Live data bypass: structured volatile data (prices, inventory, status) is queried live, not indexed at all — the architecture guidance that removes the most volatile content from the reindex problem entirely.
  3. Drift-triggered re-embed: track cumulative change; when drift crosses the threshold, schedule the full pass deliberately.
  4. Model-change migration: build the new index alongside the old, verify retrieval quality, cut over — the system never answers from half-migrated space.

The model-change migration is also a capacity moment: a full-corpus re-embed is a throughput-shaped batch job — exactly the overnight-window workload that dedicated capacity handles naturally, and that batch scheduling in orchestration layers such as the OnePlus AI Orchestration Platform queues without disturbing interactive work.

Verify: Freshness Tests, Not Assumptions

Freshness verifies by test, not by pipeline completion: a standing freshness suite asks the system questions whose correct answers depend on known-recent content and measures whether retrieval finds it — catching silent staleness (the change pipeline that quietly broke, the threshold that drifted past) — and the suite runs on a schedule, because a corpus that was fresh at launch and never tested again is one broken connector away from confidently answering from the past.

  • Known-recent questions: a maintained set whose right answers depend on documents added or changed at known times.
  • Scheduled runs: freshness is a standing test with a calendar, not a launch-day checklist.
  • Silent-failure detection: the suite fails when the incremental path quietly breaks — before users assemble the evidence for you.
  • Post-migration verification: after any model-change cutover, the suite validates the new geometry on real questions before traffic rides it.

Treat suite failures as pipeline incidents rather than quiz results: a freshness miss is a broken connector, a stalled queue, or a drifted threshold announcing itself — and the suite's value is measured in the gap between its failure and the user complaint that would otherwise carry the news.

FAQ

How often should a RAG corpus be reindexed?

On change, not on calendar: incremental indexing as content lands keeps the corpus current without scheduled ceremonies, full re-embedding waits for the documented triggers — meaningful embedding-model improvement or roughly ten to fifteen percent drift — and the calendar approach (weekly full reindexes) spends compute to fix a problem the incremental pipeline already solved.

What does an embedding-model change actually require?

A full-corpus re-embed run as a migration, not a config flip: every vector must be recomputed under the new model because the two geometries are not comparable, run as a batch embedding job sized to your corpus (the kind of overnight, throughput-shaped workload batch windows and schedulers exist for), cut over via a dual index — build new alongside old, verify, switch — so the system never answers from half-migrated space.

How do you detect RAG staleness before users complain?

With a standing freshness suite on a schedule: questions whose right answers depend on known-recent documents, asked automatically, scored on whether retrieval finds the new content — silent staleness (broken connectors, drifted thresholds) shows up as a suite failure long before it shows up as a confused user, and the gap between the two is the gap between a scheduled test and an incident report.

Previous: Private LLM Deployment: Infrastructure Requirements for Enterprise Teams
Related Articles