The National Institutes of Health’s (NIH) ambitious All of Us Research Program has reached a critical juncture in its mission to build one of the world’s most comprehensive and diverse health databases. Despite achieving a remarkable 98% consent rate among its nearly 750,000 participants for the sharing of electronic health records (EHR), a recent report by STAT reveals a significant technical bottleneck: more than 300,000 of those participants have no EHR data present in the program’s database. This discrepancy highlights a fundamental disconnect between the legal and policy frameworks governing data sharing and the technical reality of clinical data integration in the mid-2020s.
To address this "missing data" crisis, the All of Us program is shifting its strategy. The latest data release indicates a move toward "piggybacking" on existing clinical data-sharing networks—the same infrastructure hospitals use to transmit records between health systems for routine patient care. While this adjustment is viewed by experts as a necessary and pragmatic fix, it serves as a stark reminder that the "plumbing" of American healthcare data is far from a finished product. The situation reveals that while the industry has largely succeeded in building the pipes to move data, it has yet to master the art of processing the disparate, often messy information that flows through them.
The Foundation of the All of Us Research Program
Launched in 2018, the All of Us Research Program is a cornerstone of the U.S. government’s precision medicine initiative. Its primary goal is to gather health data from at least one million people living in the United States to accelerate research and improve health outcomes. Unlike previous large-scale studies that often focused on specific demographics, All of Us emphasizes diversity, seeking to include populations historically underrepresented in biomedical research.
The program collects various types of data, including physical measurements, biosamples (DNA, blood, and urine), and information from wearable devices. However, the most critical component for longitudinal research is the Electronic Health Record. EHR data provides a continuous history of a participant’s diagnoses, treatments, laboratory results, and medications. Without this data, the genomic information collected by the program loses much of its clinical context. The current gap of 300,000 records represents a significant portion of the cohort whose contributions remain partially untapped due to technical hurdles rather than a lack of participant willingness.
A Decade of Interoperability: The Policy Chronology
The challenge facing the NIH is rooted in the history of U.S. health IT policy. For the past decade, the prevailing theory of interoperability was centered on a "plumbing" metaphor. The belief was that if the federal government could force the connection of systems, data would flow seamlessly from provider to researcher.
This effort began in earnest with the 2016 passage of the 21st Century Cures Act. This landmark legislation was designed to accelerate medical product development and bring new innovations to patients faster. A key provision of the Act was the prohibition of "information blocking"—the practice of health IT developers or providers interfering with the access, exchange, or use of electronic health information.
Following the Act, the Office of the National Coordinator for Health Information Technology (ONC) and the Centers for Medicare & Medicaid Services (CMS) released a series of rules that mandated the use of standardized application programming interfaces (APIs). Specifically, the industry moved toward the Fast Healthcare Interoperability Resources (FHIR) standard. By 2024, the Trusted Exchange Framework and Common Agreement (TEFCA) had begun to take shape, aiming to establish a universal floor for interoperability across the country.
These policy efforts were largely successful in achieving their primary goal: the wires are now connected. Hospitals can query one another’s records with unprecedented ease. The 98% consent rate within the All of Us program is a testament to this success; patients believe in the system, and the legal pathways for data transfer exist. However, the 300,000-record gap proves that transmission is not the same as utility.
The "Dirty Data" Problem: Why Retrieval Fails
The reason 300,000 records remain missing from the All of Us database is not a failure of the "pipes," but a failure of the "payload." When a research program or a health system requests a patient’s full chart, what they receive is rarely a clean, structured dataset ready for immediate analysis. Instead, they often receive a digital "dump" of information that mirrors the chaotic reality of clinical documentation.
A single patient’s record might include:

- Structured Data: Discrete FHIR resources such as recent lab results or medication lists.
- Unstructured Documents: Scanned PDFs of referral letters from external specialists.
- Legacy Formats: Faxed prior-authorization forms that have been digitized but not OCR-processed (Optical Character Recognition).
- Inconsistent Field Notes: Visit notes exported as PDFs with headers and fields that vary between EHR vendors like Epic, Cerner, or Meditech.
In many cases, a patient’s history spans decades, including records from systems that were never designed to be machine-readable. When these files arrive at a research database, they are often rejected or placed in a "pending" queue because they do not meet the strict formatting requirements of the research repository. The pipe worked—the data moved from the hospital to the NIH—but because the data wasn’t "clean," it effectively doesn’t exist for the purposes of the study.
The Shift Toward Advanced Data Engineering
The current state of the industry suggests that the next phase of interoperability will be a data-engineering problem rather than a policy one. Experts argue that the tools required to bridge this gap have matured significantly over the last two years, driven largely by advancements in artificial intelligence and machine learning.
The emerging solution is a hybrid model that combines probabilistic extraction with deterministic validation.
- Probabilistic Extraction: This involves AI systems, such as Large Language Models (LLMs) and specialized medical NLP (Natural Language Processing) tools, that can read scanned, handwritten, or inconsistently formatted documents. These systems "read" a document much like a human would, identifying key clinical concepts even when they are buried in a narrative note or a poorly scanned fax.
- Deterministic Validation: Because AI can be prone to "hallucinations" or inaccuracies, a second layer of rules-based logic is applied. This layer checks every extracted field against medical ontologies (like SNOMED or ICD-10) and pre-defined clinical bounds. If a system extracts a blood pressure reading that is physiologically impossible, the deterministic layer flags it for human review.
This pairing is critical because research and clinical use demand a level of accuracy that probabilistic AI cannot yet achieve on its own. Conversely, traditional rules-based systems are too rigid to handle the sheer diversity of real-world medical records. By combining the two, the time required to integrate a new health system’s legacy data into a research database has dropped from years of custom coding to mere weeks.
Implications for Precision Medicine and Clinical Research
The inability to process unstructured data has profound implications for the future of precision medicine. If 40% of a research cohort’s data is missing because it exists in PDF format rather than structured FHIR resources, the conclusions drawn from that research may be skewed.
For the All of Us program, the stakes are particularly high. The program aims to identify how environment, lifestyle, and biology interact to influence health. If the data gap disproportionately affects participants from smaller, rural hospitals or older clinics—facilities more likely to rely on legacy systems and scanned documents—the research could inadvertently exclude the very "diverse" populations it was designed to study.
Furthermore, the automation of data extraction offers a more rigorous evidentiary record than manual chart abstraction. Historically, researchers would hire "clinical abstractors" to manually read charts and enter data into databases. This process is slow, expensive, and prone to human error. Modern automated systems create a digital audit trail, showing exactly which document a piece of data came from and the confidence level of the extraction, providing a level of transparency that manual processes cannot match.
Future Outlook: Beyond the 300,000-Record Gap
The decision by All of Us to leverage existing clinical data-sharing networks is a tactical victory that will likely close a significant portion of the 300,000-record gap. By using the networks that hospitals already trust and use daily, the program reduces the friction of data acquisition.
However, this move is only the first step. As more records flow into the program, the "messy remainder"—the scans, faxes, and legacy exports—will become the primary bottleneck. The success of the next decade of healthcare research will depend on whether the industry treats data engineering with the same level of urgency it gave to data transport.
The lesson from the STAT report is that interoperability, as defined by the ability to move a record from System A to System B, has largely succeeded. The new challenge is ensuring that what arrives at System B is actually usable. For research programs, health systems, and payers, the shift from "connecting pipes" to "structuring content" represents the final frontier of the digital health revolution. Those who embrace advanced engineering solutions to transform unstructured documents into validated data will be the ones to finally close the gap between patient consent and scientific discovery.
