Rescue · Technology
AI EHR parsing won't save your under-enrolling study. It will just reach the same ineligible patients faster.
If a study is failing because the underlying patient data is out of date, speed is not the remedy. Speed is the multiplier on the mistake.
Rock Enroll editorial · September 2026
The industry has repeated the same number for a decade: the large majority of trials — commonly cited between 60% and 80% depending on the sample and the definition — fail to meet their enrollment timelines. That number is quoted so often that it has stopped being a question. It should be one. If enrollment fails routinely, at that rate, across sponsors, indications and vendors, then it is not a series of unlucky studies. It is a property of the method everyone is using.
The pre-existing condition was the database
Ask what a recruitment campaign actually contacted, and in most cases the answer is a list. A patient database, a registry, a claims-derived cohort, an EHR extract, a panel someone bought. The same lists circulate between vendors, and the same people on them are contacted for study after study.
Two things degrade that list, and they degrade it continuously.
1. Contactability decays
A record collected two years ago carries a phone number, an email address and a consent state from two years ago. People change numbers, abandon inboxes, move, and stop answering unknown callers — particularly after they have been called about three previous studies they did not qualify for. Nothing about the record announces that it has gone stale. It looks identical to a fresh one right up until nobody answers.
2. Eligibility decays faster
This is the part that gets skipped. A patient's eligibility is a snapshot of their health at the moment the data was captured. Between then and now, their HbA1c moved. They started a medication that is now an exclusion. Their disease progressed past the stage the protocol accepts, or improved out of it. They had the surgery. They enrolled in someone else's trial. The record still says “matches criteria.” The person no longer does.
So a database has two independent decay curves running against it, and the second one is invisible to the system holding the data. Nobody re-measured the patient. The database has no idea.
What AI actually changes
Put a language model on top of EHR notes and it will do genuinely impressive work: read unstructured narrative, resolve messy phrasing into structured criteria, and identify candidate matches at a scale no coordinator could. This is real capability, and it is worth being precise about what it is capability at.
It is capability at querying the record you already have. It parses faster, it parses more, it parses text humans would have skipped. What it cannot do is tell you whether the record is still true. An AI matching engine reading a stale chart produces a confident, well-structured, fully-reasoned match to a patient who became ineligible eleven months ago. The output is cleaner than a human's. It is wrong in exactly the same way.
That is the whole argument, and it is not an argument against AI. Reaching the same ineligible participants faster is the only difference AI makes when the substrate is outdated. You fail faster, at higher confidence, with better dashboards.
The cost of refreshing the substrate
The honest fix — re-contacting a cohort, re-measuring current health status, re-consenting, re-verifying contact details, and doing it on a rolling basis so the data never ages past usefulness — is not a software problem. It is an expensive, continuous, operational one. Sponsors who have the millions and the years to maintain a live cohort get real value from AI on top of it, because the layer underneath is true. Nearly everyone else is applying modern inference to a dead snapshot.
The measurement problem underneath it
There is a second reason database recruitment persists despite the failure rate: it is very difficult to price. When a vendor works a list, what did an interested, eligible, contactable patient cost? Not the cost per record. Not the cost per outreach attempt. The cost of the one person who was genuinely reachable, genuinely eligible today, and genuinely willing to show up. Most sponsors cannot answer that, because the denominator — how many of those records were already dead — is unknowable from the inside.
Advertising has one structural property that list-working does not: the demand is expressed now. A person responding to an ad this week is contactable this week and interested this week, by definition. Their eligibility still has to be screened, and plenty will fail. But the cost of reaching genuine, current interest is directly observable — spend divided by the people who actually raised a hand — and it can be measured per channel, per creative, per geography, per site catchment, and adjusted while the study is running.
That is not a claim that advertising is cheap or that it fixes every study. It is a claim about what you can know. A channel with an observable cost per genuine interest can be managed. A database with an unobservable decay rate can only be hoped at, and hope is what the 60–80% figure is measuring.
Questions worth asking a vendor pitching AI-driven EHR recruitment
- When was each record in the matched cohort last updated — not last accessed, last clinically updated?
- What share of your AI-identified matches were reachable on the first contact attempt? What share on any attempt?
- Of those reached, what share were still eligible when re-screened against current health status — and how does that compare with what the parse predicted?
- How many of these patients were contacted for another study in the last twelve months?
- What is the cost per randomized patient from this cohort, and how does it compare with a channel where interest is expressed in real time?
A vendor who can answer these has a live cohort and their AI is doing real work. A vendor who cannot is selling faster access to an archive, and the pilot will look excellent for about six weeks. If you are at the point of choosing between vendors, our directory of enrollment and recruitment companies sets out the criteria we use to assess them.
The diagnostic sequence for a study already behind is the same either way: find the failing stage before approving more spend. The four-stage funnel diagnosis covers how to do that.
Disclosure
Rock Enroll is published by CT Scan, Inc., the company behind DYNO . That is our commercial interest and you should read this piece with it in mind. The questions listed here are ones you can put to any vendor, including us.
More in this series
All rescue guides