Analysis · Literature review

How old is the data behind an AI-matched patient cohort?

Every AI matching claim is a claim about a database. The published work evaluates the matching. It rarely evaluates the age of what is being matched.

Rock Enroll editorial · September 2026

The trial-matching literature has moved fast. Large language models are now being evaluated against real electronic health records for eligibility prescreening, with published work reporting genuine gains in accuracy and staff time over manual chart abstraction — including a randomized evaluation of human-AI teaming for oncology prescreening in Nature Communications, a locally-deployed adaptation of TrialGPT evaluated in JAMIA, a multimodal matching pipeline validated in Communications Medicine, and AI-enabled chart review for cardiomyopathy trial eligibility in the Journal of Cardiac Failure.

Read the designs and a pattern appears. These studies are overwhelmingly retrospective evaluations against the record: the model's judgement is compared with what a trained human abstractor concludes from the same chart. That is the right way to measure a parser. It is not a measurement of whether the patient is eligible today, reachable today, or willing today — because the ground truth is the chart, and the chart is the thing whose age is in question.

The distinction this piece is drawing: parsing accuracy is agreement between the model and the record. Cohort validity is agreement between the record and the person. A system can be excellent at the first and produce an unusable list because of the second, and almost no published evaluation reports both.

What is known about the substrate

Contact details decay measurably

The most concrete evidence comes from follow-up research rather than recruitment research. A cross-sectional study of orthopedic patients in Clinical Trials examined how patient contact information changes over time specifically because of what it does to research validity, and cancer-screening trials have published dedicated tracking-and-tracing methods for the participants their records could no longer reach. The existence of a tracing literature is itself the finding: cohorts that were consented, engaged and actively followed still lose contactability fast enough to require a discipline for recovering it. An EHR extract nobody has contacted in three years has no such discipline behind it.

Consent-to-recontact is not the same as reachable

Registries built specifically for recruitment — where patients have explicitly agreed to be contacted about future studies — still publish comparative evaluations of contact strategies because response rates are the limiting factor, and rare-disease contact registries have needed automated communication systems to work at all. These are the best-case databases in the field: opted-in, motivated, maintained. If reach rates are the headline problem there, the base rate for a purchased or claims-derived list is not a mystery.

Eligibility drift is the part nobody measures

Contactability at least announces itself — the call fails and you know. Eligibility drift does not. Between the date a record was captured and the date it is parsed, a patient may have started a medication that is now an exclusion, progressed past the accepted disease stage, improved out of it, had surgery, or joined another study. The record still reads as a match.

The last two years make this sharper rather than academic. Widespread adoption of GLP-1 receptor agonists moves weight, HbA1c, blood pressure and lipid profiles across whole populations of patients — precisely the metrics that gate eligibility in metabolic, cardiovascular, hepatic and renal protocols. A record captured before that patient started therapy describes someone who no longer exists. Every therapeutic class that enters routine practice does this to some slice of the historical record, silently and continuously.

The measurement gap, stated plainly

We could not find published work reporting the metric that would settle the question for a sponsor buying AI-driven recruitment. That metric is straightforward to define:

  • Of patients an AI system flagged as eligible from stored records, what share were reached on any contact attempt?
  • Of those reached, what share were still eligible when re-screened against current clinical status?
  • Stratified by the age of the underlying record, in months.

Until vendors publish that third stratification, an accuracy figure quoted for a matching engine is a statement about its reading comprehension, not about the cohort it hands you. If you are evaluating one, ask for it directly — the vendor assessment criteria in our directory include it, and the full argument for why it is the decisive question is in why AI EHR parsing won't fix an under-enrolling study.

Nothing here argues against AI in recruitment. It argues that the published evidence supports a narrower claim than the one being sold: these systems read records well. Whether the records are still true is a separate question, and it is currently unevidenced.

Disclosure

Rock Enroll is published by CT Scan, Inc., the company behind DYNO. This analysis uses only public records; the method is stated so anyone can reproduce or contradict it.