Beyond the Form: Reimagining Clinical Trials
A Phase III protocol now collects around 5.9 million datapoints, and it can still be hard to say why a participant stopped treatment. The scarce asset is not data, it is context. A systems view of when clinical trials impose structure, and what that decision costs.
Capture richly. Preserve context. Structure intelligently. Validate rigorously.
A question increasingly put to clinical-development teams is: Can AI accelerate clinical trials?
It is a reasonable question. It may also be the smaller one.
If we were designing clinical trials today, with AI available from the beginning, would we design them in the same way?
Would the electronic case report form still be the primary interface through which we decide what becomes useful clinical-trial data? Would patients, investigators and sites still need to translate complex clinical experience into predefined variables as early as possible?
And as machine agents begin to participate in clinical workflows alongside humans, how should we decide which actor observes, interprets, decides, verifies and acts?
These are not primarily questions about AI capability. They are questions about systems design and information architecture. And systems thinking may be a useful place to start.
We collect more data than ever. But are we preserving enough context?
Start with a paradox. Across 105 benchmarked protocols, contributed by fourteen of the fifteen sponsor companies in a Tufts CSDD and TransCelerate working group, Getz and colleagues reported that a Phase III protocol now collects a mean of approximately 5.9 million datapoints, increasing by roughly 11% annually since 2020. Nearly one-third of Phase III procedures and their associated datapoints were classified as non-core or non-essential. Yet 74% of data associated with non-core procedures in Phase II and III ultimately appeared somewhere in the Clinical Study Report. [1]
One interpretation is that trials simply collect too much data. There is probably truth in that. But it does not explain everything. If data volume were the only problem, reducing data volume should solve it.
A more interesting question is: With millions of datapoints available, why can it still be difficult to answer some of the most human questions about what happened during a trial?
Why did this participant actually discontinue? Why was another participant hospitalised despite intensive monitoring? What changed between visits? What did the participant tell the investigator that was clinically meaningful but never became an analysis variable?
Data volume is not information richness. Information richness is not clinical understanding. Perhaps the scarce asset is not data. Perhaps it is context.
Start with the system
A clinical trial is not a collection of forms. It is an interconnected socio-technical system. Patients, caregivers, investigators, study coordinators, sponsors, CROs, laboratories, Data Management, medical monitors, safety teams, statisticians, technology platforms and regulators all interact. Information moves among them. So do decisions. Every interface between those actors is also an information transformation.
At a site visit, an investigator may hear a complex account from a participant. Part of it becomes a medical history entry. Part becomes an adverse event. Part becomes a medication change. Part becomes free text. Part remains in source documentation. Part may remain only in the understanding of the people who were present.
At each handoff, the system decides, often implicitly, what survives.
One account, six destinations. Three of them reach analysis, two are kept without being analysable, and the most explanatory part is not recorded at all.
Systems thinking invites us to ask different questions: Where is context created? Where is it transformed? Where is it compressed? Where can it disappear? Which later decisions depend on what was lost?
Viewed this way, the eCRF is not the problem. It is one important interface in a much larger system.
We decide much of what matters before the patient arrives
Clinical trials necessarily anticipate the future. Before first patient in, a familiar chain is established: protocol, endpoints, assessments, eCRF, variables, controlled terminology, analysis. That architecture delivers enormous value. It provides consistency. It allows aggregation. It supports reproducibility. It makes statistical analysis possible. It enables regulatory review. But it also has an unavoidable consequence: the structure exists before the individual participant’s experience does. Whatever was not anticipated may still survive in source records or free-text fields, but it does not necessarily become accessible downstream as analysable information.
This resembles a schema-on-write architecture: the structure is already decided, before information enters the system.
Historically, that was not simply sensible. It was necessary. Computers processed predefined fields far more effectively than human language.
The important question is therefore not whether the eCRF was the right solution. It was. The question is whether the technical conditions that made it the only reasonable conceptual starting point still hold.
Structure represents reality, but does not contain all of it
The difference becomes clearer when we examine narrative clinical information.
Seinen and colleagues compared structured clinical codes with concepts extracted from free-text notes across a Dutch primary-care database containing 1.8 million patients and 14 million visits. Only 13% of concepts identified in free text had a structured counterpart in the record; at the individual-visit level, the figure was 7%. In the other direction, 42% of structured concepts appeared in the narrative. [2]
These numbers should not be transferred directly to clinical trials. The study involved Dutch routine-care records, not GCP trial data, and the result depends on the quality of concept extraction and the similarity threshold used. But the direction of the finding matters. Narrative and structure are not interchangeable representations of clinical reality.
Consider a participant who stops treatment. The database eventually contains: Study treatment discontinued: Yes. Reason: Adverse event. Technically correct. But perhaps fatigue had increased over several weeks. Working had become difficult. The participant’s husband could no longer drive her to study visits. She was concerned that worsening symptoms meant the treatment was failing. Several of these issues had been discussed during site conversations. “Adverse event” is not wrong. It is thin. What disappears is not necessarily accuracy. It is resolution.
Was this primarily a tolerability problem? A logistical problem created by study design? An expectation-management problem? Several interacting causes? The database can tell us what happened. Context may be needed to understand why.
That context is not necessarily unrecoverable. Suvalov and colleagues applied large language models to free-text anamneses from a 10% sample of the Estonian population and asked two questions of the narrative: why was the medication stopped, and who stopped it. Reasons were extracted with a precision of 0.93 to 0.98 and classified into a clinician-developed taxonomy with weighted F1 scores of 0.81 to 0.84. Adverse reactions accounted for 70% of statin discontinuations. Identifying who initiated the decision, the patient or the physician, proved harder, at weighted F1 scores of 0.64 to 0.78. [4]
This was routine care rather than a regulated trial. Two observations still carry over. The narrative held information the structured record did not, and the more interpretive the question, the more verification its answer requires.
Context should become a first-class information asset
This does not mean recording everything indiscriminately. “More data” should not become the objective. Clinical trials already have enough problems with unnecessary data collection.
The objective should be more precise: preserve clinically meaningful context where its future value cannot be fully anticipated at the moment of capture. That is a different design principle. And AI changes what is technically possible.
Humans communicate context naturally through language. Yet research systems traditionally ask humans to communicate with computers primarily through forms, checkboxes, dropdowns, fixed questionnaires and predefined fields. That reflected the capabilities of machines, not the natural information-processing capabilities of people. Large language models change part of that constraint.
Recent work indicates how far that now reaches. Dickerson and colleagues gave off-the-shelf models unnormalised, unlabelled clinical notes, pathology reports and medication records from a breast-cancer cohort, and asked them to abstract a longitudinal dataset. The strongest model reached 99% concordance for recurrence status, 100% for germline BRCA1/2 pathogenic variants and 96% for HER2 status, while its extraction of systemic therapies approached the variability observed between oncologists. Survival estimates derived from the model-abstracted dataset closely matched those derived from the expert-abstracted one. The work is a preprint and concerns retrospective oncology records rather than a regulated trial, so it demonstrates feasibility rather than a validated trial workflow. [5]
It becomes conceivable to preserve a richer interaction first and then use computational systems to identify concepts, propose structure, locate supporting evidence and make the result available to downstream processes.
The preserved context in the middle is the architectural difference. It is what allows a question nobody anticipated to be asked later.
The distinction is important. This is not voice, then AI, then database. The preserved context in the middle is the architectural difference.
Verification, in turn, need not mean that a human repeats every step. Leinonen and colleagues found that misclassified findings in Finnish records carried markedly higher model uncertainty than correct ones, so that reviewing a small share of cases captured all observed errors. Uncertainty can itself become a routing signal, concentrating human judgement where it changes the outcome. This is likewise a preprint, and outside the trial setting. [6]
If an extraction is later found to be incomplete, the source remains available. The extraction specification can change. Another model can interrogate it. A human can review it. A question that nobody anticipated at study design may potentially be asked later. Preserving source context creates optionality.
Reimagine when we impose structure, not whether structure is needed
This is not an argument against structured clinical data. Statisticians need structured datasets. Regulators need reproducibility. Standards such as CDISC remain essential. Endpoints must still be prespecified. Laboratory values, dosing, dates and other deterministic information belong in appropriate structured systems. Validated patient-reported outcome instruments depend on controlled wording, order, recall periods and response options; they cannot simply be converted into adaptive conversation without evidence that measurement properties have been preserved. [7]
The question is narrower: Does every clinically useful observation have to reach its final structure at the moment it is first captured?
The dominant approach is approximately: structure first, then capture, then analyse.
For some information, an alternative may be: capture richly, preserve, structure for purpose, validate, then analyse.
This resembles schema-on-read: the underlying record is retained, while fit-for-purpose structure can be derived when the question is asked.
| Design dimension | Schema-on-write, the eCRF today | Schema-on-read, preserved context |
|---|---|---|
| When structure is decided | Before the first participant, at protocol design | When the question is asked, after capture |
| What is stored | The predefined variable | The source record, and the variable derived from it |
| Optimised for | Consistency, aggregation, reproducibility, regulatory review | Resolution, re-interrogation, questions nobody anticipated |
| Failure mode | Whatever was not anticipated is compressed away and does not come back | Interpretation has to be verified and provenance has to be carried |
| Changing your mind | Needs an amendment, and retrospectively it is often impossible | Run the extraction again against the same source |
| Fits | Endpoints, laboratory values, dosing, dates, validated instruments | Reasons, trajectories, burden, the context around an event |
The practical model is therefore hybrid. Not structured data or narrative. Not eCRFs or conversations.
But: preserve the structure that must remain explicit while preserving richer context where premature compression may destroy future information value.
In practice, that line can be drawn quite concretely.
| Information | Where it belongs | Why |
|---|---|---|
| Laboratory values, vital signs, dosing, dates | Structured at capture | Deterministic and machine-generated. There is no interpretation to preserve. |
| Endpoints and their assessments | Structured and prespecified | Statistical analysis and regulatory review depend on it. |
| Validated PRO instruments | Structured, wording unchanged | Measurement properties depend on wording, order, recall period and response options [7] |
| Adverse events | Hybrid | The coded term stays explicit. The account behind it is worth keeping. |
| Reason for discontinuation | Context first, structure derived | ”Adverse event” is accurate and thin. The real reason is usually several at once [4] |
| What changed between visits | Context preserved | Rarely anticipatable as a variable, and often the clinically interesting part. |
| Travel, caregiver, work, burden | Context preserved | Almost never a variable, yet it drives retention and protocol feasibility. |
Beyond the form
For decades, clinical trials required humans to communicate in ways computers could understand. Forms. Fields. Codes. Checkboxes. That architecture brought standardisation, reproducibility and rigor. It should not be casually discarded.
But for the first time, machines can increasingly interpret the way humans naturally communicate. That means the design space has changed. The question is no longer simply: How can AI fill the existing form faster?
It is: What should the clinical-trial system look like when both humans and machines can interpret information, make propositions and participate in workflows?
This line of thinking is not ours alone. At the Mayo Clinic, Al Zahidy and colleagues are assembling exactly the asset such a system would require. Their observational protocol records real patient-clinician encounters with 360-degree video and dual-channel audio, links each encounter to a post-visit survey and to the electronic health record, and treats the result as a foundational dataset for downstream AI research. Their reasoning is close to the argument made here: models trained on electronic health records learn biological measures, but rarely the interaction in which care is understood, negotiated and delivered, and a model trained only on that record inherits a narrow biomedical view of medicine. The early feasibility figures suggest the approach is workable in a real clinic. 35 of 36 eligible clinicians and 212 of 281 approached patients consented, 76% of consented encounters produced a complete recording, and 96% of post-visit surveys were returned. [3]
Two things follow from that. Other groups are already working to distil more from patient-clinician communication than checkboxes and tables. And the models themselves are being developed towards these richer information formats, rather than only towards the fields we have historically been able to store.
Systems thinking suggests several principles: Preserve context where premature compression destroys value. Keep explicit structure where science and regulation require it. And most importantly: Capture richly. Preserve context. Structure intelligently. Validate rigorously.
If we were designing clinical trials from scratch today, with AI available from day one and human and machine agents both participating in the system, would the eCRF still be where the patient’s data journey begins?
Perhaps reimagining the clinical trial starts by looking beyond the form.
Sources
- Getz K, Botto E, Calduch Arques A, et al. Insights Informing Strategies for Optimizing the Collection of Clinical Trial Data. Therapeutic Innovation & Regulatory Science. 2026;60(2):563-574. doi:10.1007/s43441-025-00899-4. link.springer.com
- Seinen TM, Kors JA, van Mulligen EM, Rijnbeek PR. Using Structured Codes and Free-Text Notes to Measure Information Complementarity in Electronic Health Records: Feasibility and Validation Study. Journal of Medical Internet Research. 2025;27:e66910. doi:10.2196/66910. jmir.org
- Al Zahidy M, Guevara Maldonado K, Vilatuna Andrango L, et al. Longitudinal and Multimodal Recording System to Capture Real-World Patient-Clinician Conversations for AI and Encounter Research: Protocol for an Observational Study. JMIR Research Protocols. 2026;15:e84688. doi:10.2196/84688. researchprotocols.org
- Suvalov H, Umov N, Malk M, et al. Extracting and Classifying Drug Discontinuations From Estonian Electronic Health Records: Development and Validation Study. Journal of Medical Internet Research. 2026;28:e86183. doi:10.2196/86183. jmir.org
- Dickerson JC, McClure MB, Shaw M, et al. Fully Automated Abstraction of Longitudinal Breast Oncology Records with Off-The-Shelf Large Language Models. medRxiv. 2026. Preprint. doi:10.64898/2026.03.23.26349012. medrxiv.org
- Leinonen J, Knuuttila J, Pamilo S, Kurki S, Koskinen M. Uncertainty-aware extraction of clinical findings from Finnish EHRs using open large language models. medRxiv. 2026. Preprint. doi:10.64898/2026.07.07.26355248. medrxiv.org
- U.S. Food and Drug Administration. Guidance for Industry. Patient-Reported Outcome Measures: Use in Medical Product Development to Support Labeling Claims. December 2009.
Rethinking where your trial data begins
We help clinical development teams decide where structure belongs, where context is worth preserving, and how AI fits into the workflow around both. Let us talk about your setting.
Frequently asked questions
Does this mean abandoning the eCRF?
No. The eCRF was the right answer to the technical conditions of its time, and much of what it does remains necessary. Statisticians need structured datasets, regulators need reproducibility, endpoints must still be prespecified. The narrower question is whether every clinically useful observation has to reach its final structure at the very moment it is first captured.
What does schema-on-read mean in a clinical trial?
It is a term borrowed from data architecture. Schema-on-write decides the structure as information enters the system, which is what an eCRF does. Schema-on-read keeps the underlying record and derives fit-for-purpose structure when a question is actually asked. In a trial the practical model is hybrid: explicit structure where science and regulation require it, preserved context where premature compression would destroy future value.
How would a sponsor start without redesigning everything?
Pick one narrow question whose answer is currently thin. Reasons for treatment discontinuation are a good candidate, because the structured field is almost always accurate and almost always insufficient. Preserve the context around that one question, derive structure from it, validate the result against what the site actually documented, and compare the two accounts.
Where does AI stop and human judgement begin?
At the point where the consequence of being wrong outweighs the cost of a review. Capability is not authority: a model that can identify an adverse-event concept is not thereby permitted to decide seriousness or relatedness. Useful systems make that boundary explicit, and route uncertain or consequential outputs to a human rather than asking humans to re-perform every task.