A RAG system is only as current as its corpus, and corpora rot quietly: documents change, connectors break, thresholds drift, and the system keeps answering — confidently, from the past. Most teams discover staleness through a user complaint about an answer everyone knows is outdated. The alternative is an operated refresh pipeline: triggers that are written down, incremental mechanics that avoid brute force, and tests that catch silence before users do.
Prerequisites: Know Your Two Triggers
Two triggers govern the whole pipeline: content drift — the documented threshold sits around ten to fifteen percent of the corpus changed, before which incremental indexing handles everything and after which re-embedding pays — and embedding-model change, the non-negotiable event where every vector must be recomputed because embeddings from different models are not comparable and a half-migrated index answers from two incompatible geometries at once.
| Trigger | Threshold | What it forces |
| Content drift | ~10-15% of corpus changed | Full re-embed pays; below it, incremental suffices |
| Embedding-model change | Any switch, any time | Mandatory full re-embed — vectors are not cross-compatible |
| Ordinary change flow | Continuous | Incremental indexing on arrival; no ceremony |

The model-change trigger deserves its non-negotiable status in plain terms: with the same embedding model, re-indexing is unnecessary; the moment the model changes, everything must be re-embedded — because similarity search across two embedding spaces is arithmetic without meaning, and the failure mode is retrieval that looks fine and answers wrong.
Build the Refresh Pipeline: Incremental, Not Brute Force
Build incremental, not brute force: changes flow through an incremental indexing path as they land (new and updated chunks re-embedded on arrival, structured live data queried directly rather than indexed at all), full re-embedding reserved for the two triggers — and the pipeline treats the model-change re-embed as a migration project with dual-index cutover, because the expensive periodic full reindex that brute-forces freshness is the documented anti-pattern the cheaper design exists to replace.
- Incremental on change: changed documents flow to re-embedding on arrival — the corpus stays current without scheduled ceremonies.
- Live data bypass: structured volatile data (prices, inventory, status) is queried live, not indexed at all — the architecture guidance that removes the most volatile content from the reindex problem entirely.
- Drift-triggered re-embed: track cumulative change; when drift crosses the threshold, schedule the full pass deliberately.
- Model-change migration: build the new index alongside the old, verify retrieval quality, cut over — the system never answers from half-migrated space.
The model-change migration is also a capacity moment: a full-corpus re-embed is a throughput-shaped batch job — exactly the overnight-window workload that dedicated capacity handles naturally, and that batch scheduling in orchestration layers such as the OnePlus AI Orchestration Platform queues without disturbing interactive work.
Verify: Freshness Tests, Not Assumptions
Freshness verifies by test, not by pipeline completion: a standing freshness suite asks the system questions whose correct answers depend on known-recent content and measures whether retrieval finds it — catching silent staleness (the change pipeline that quietly broke, the threshold that drifted past) — and the suite runs on a schedule, because a corpus that was fresh at launch and never tested again is one broken connector away from confidently answering from the past.
- Known-recent questions: a maintained set whose right answers depend on documents added or changed at known times.
- Scheduled runs: freshness is a standing test with a calendar, not a launch-day checklist.
- Silent-failure detection: the suite fails when the incremental path quietly breaks — before users assemble the evidence for you.
- Post-migration verification: after any model-change cutover, the suite validates the new geometry on real questions before traffic rides it.
Treat suite failures as pipeline incidents rather than quiz results: a freshness miss is a broken connector, a stalled queue, or a drifted threshold announcing itself — and the suite's value is measured in the gap between its failure and the user complaint that would otherwise carry the news.
FAQ
How often should a RAG corpus be reindexed?
On change, not on calendar: incremental indexing as content lands keeps the corpus current without scheduled ceremonies, full re-embedding waits for the documented triggers — meaningful embedding-model improvement or roughly ten to fifteen percent drift — and the calendar approach (weekly full reindexes) spends compute to fix a problem the incremental pipeline already solved.
What does an embedding-model change actually require?
A full-corpus re-embed run as a migration, not a config flip: every vector must be recomputed under the new model because the two geometries are not comparable, run as a batch embedding job sized to your corpus (the kind of overnight, throughput-shaped workload batch windows and schedulers exist for), cut over via a dual index — build new alongside old, verify, switch — so the system never answers from half-migrated space.
How do you detect RAG staleness before users complain?
With a standing freshness suite on a schedule: questions whose right answers depend on known-recent documents, asked automatically, scored on whether retrieval finds the new content — silent staleness (broken connectors, drifted thresholds) shows up as a suite failure long before it shows up as a confused user, and the gap between the two is the gap between a scheduled test and an incident report.