A Reference Framework for AI-Native eDiscovery

Discovery, built for data
that never has to move.

The EDRM* was built in 2005 for a world where data had to move before anyone could look at it. This model is built for a world where a governed LLM can query a corporation's live systems in place. Five moving stages, one constant foundation.

The Core Shift

Nine phases existed because data had to sit still.

Identification, collection, processing, review, analysis, production, presentation, each was a discrete event because a document had to be physically copied out of a live system before a human or a keyword search could touch it. Take that constraint away, and the phases built on top of it collapse into something faster, and something new has to take their place.

The Living Discovery Model flowchart: Connected Corpus, Continuous Indexing, Retrieval-Augmented Discovery, AI-First Review with Validation, and Explainable Production, built on a Preservation foundation
The Living Discovery Model — five stages on one foundation, not nine sequential phases.
Five Stages

What changes, and what it replaces

01
Connected Corpus
was Identification + Collection
Permissioned access into live systems. Data is queried where it lives, nothing copied out.
02
Continuous Indexing
was Processing
Chunking, embedding, and deduplication run always, not as a batch job before review starts.
03
Retrieval-Augmented Discovery
was Identification + Analysis
Conversational, iterative querying. Finding and evaluating happen in the same turn.
04
AI-First Review with Validation
was Review
The model makes the first pass. Humans validate through statistical sampling.
05
Explainable Production
was Production + Presentation
Every document ships with its reasoning trail: what surfaced it, and why.

Preservation

The foundation underneath all five stages, not a phase in the sequence. Data is held in place and retained. An immutable, auditable log of what the system accessed, when, and what it did is added. This replaces the old defensibility anchor, the physical act of collection, which no longer exists once data is never moved.

The Open Question
Courts will ask how you know you found everything, and what you didn't find, and why.

Today's recall and precision statistics were built for a search that runs once against a static index. They don't map cleanly onto a conversational, iterative, non-deterministic retrieval process. Whoever solves validation for AI-native discovery first sets the standard the rest of the industry gets measured against.

* EDRM.net