Every RWE study that uses EHR data eventually has to answer this question: are the patients identified by our code criteria actually the patients we want? It sounds simple. In practice, the gap between what a billing code describes and what a clinical phenotype requires is one of the most consistent sources of study quality problems we see.
This piece is specifically about ICD-10 codes and phenotyping, because that is where the friction is most acute. Other structured data fields, lab values, medication orders, procedure codes, have their own challenges. But diagnosis codes are what most cohort definitions start with, and they are the most frequently misused. Understanding why requires understanding what ICD codes actually are and what they were designed to do.
What ICD-10 Codes Capture (and What They Do Not)
ICD-10 codes are an administrative classification system originally developed for mortality statistics and adapted for hospital billing. A code like I50.22 (chronic systolic heart failure, acute on chronic) is assigned by a coder reviewing the provider's documentation after an encounter, primarily to support billing. The code is an abstraction of the clinical record, not a primary source of clinical information.
This matters because the clinical intent and the billing intent are not identical. A provider who documents "history of heart failure, currently compensated" creates a record where a coder might assign an I50.x code for accurate comorbidity capture, even though the condition is not driving the current encounter. A different coder at a different institution might not assign that code. The same clinical reality can produce different coding patterns depending on coder training, facility coding guidelines, and payer-specific rules.
The published literature on code-based phenotyping shows a wide range of positive predictive values. For some conditions with strong code-outcome correlation, like hip fracture or myocardial infarction with distinctive ICD codes, code-only definitions can achieve PPV above 85% against chart review. For conditions with heterogeneous documentation patterns, like depression, COPD, or early-stage chronic kidney disease, code-only definitions can fall to 50-60% PPV. The patients included in those cohorts include a substantial proportion who do not have the condition as clinically intended by the study.
What a Clinical Phenotype Actually Requires
A phenotype is a computable patient characterization that attempts to capture a clinically meaningful condition as it would be defined by a clinician. The PhEWAS (phenome-wide association study) literature and the PCORnet phenotype library have done significant work to document validated phenotypes for common conditions, and the consistent finding is that high-PPV phenotypes almost always require multiple evidence sources.
A validated phenotype for type 2 diabetes, for example, might require two ICD-10 E11.x codes from separate encounters, plus evidence of diabetes-appropriate medication (metformin, GLP-1 agonists, insulin) or a glycated hemoglobin result above a specified threshold. That combination of structured data elements achieves higher PPV than any single element alone, because the combination reduces the chance that any one element reflects a billing artifact rather than a clinical reality.
The problem is that even multi-element structured phenotypes have a ceiling. For conditions where the clinically meaningful discrimination depends on information that is documented in notes but not codified in structured fields, no combination of billing codes and lab thresholds can get there. Systolic versus diastolic heart failure subtypes, NYHA functional class, oncology stage at diagnosis, and cognitive impairment severity are all examples of clinically important distinctions that require note text to capture reliably.
Where the Bridging Happens in Practice
The transition from code-based to phenotype-based cohort definition happens at a specific point in most RWE workflows: when the study team presents their cohort definition for internal scientific review and is asked to provide evidence that the cohort contains the intended patient population. At that point, there are roughly three paths.
The first path is to conduct a manual validation chart review, sampling a subset of patients and confirming against the clinical note that they meet the phenotype criteria. This is the gold standard but is expensive and time-limited; it also applies only at a point in time and does not scale to larger cohort refreshes.
The second path is to accept the code-based definition as a reasonable approximation, cite the published literature on code performance for that condition, and acknowledge the misclassification as a limitation. Many observational studies take this path, particularly for conditions with published high-PPV estimates. The limitation section covers it, and reviewers accept it as a field-standard approach.
The third path is to use automated extraction from clinical notes to augment the structured data. Instead of relying on ICD codes alone, you extract the relevant clinical findings from the notes of your candidate population and use the extracted evidence to confirm or exclude patients. This approach is more expensive to set up than code-only queries, but cheaper than full manual chart review for large cohorts, and it produces a more defensible cohort definition for outcomes where the clinical literature does not provide strong code-performance estimates.
A Concrete Example: Distinguishing CKD Stages
Chronic kidney disease is a useful case because the clinically important distinctions among stages matter a great deal for cohort selection and are inadequately captured in ICD codes. ICD-10 N18.1 through N18.6 represent stages 1 through 5 and end-stage renal disease. In theory, a code-based query could identify patients at a specific CKD stage. In practice, stage assignment in the codes often reflects documentation from one encounter without being updated as GFR changes over time. The same patient might have N18.3 in their problem list from a 2022 encounter while their most recent calculated GFR documents stage 4 progression.
A phenotype that requires current GFR within a study window, confirmed against a threshold appropriate for the target stage, combined with a clinical note that confirms the CKD diagnosis and documents the staging rationale, is more accurate than either code alone or GFR alone. The note adds the clinical context that the structured values lack: is the GFR trajectory declining, stable, or a one-time reading in a post-acute context where the result does not represent the patient's baseline?
We ran an internal comparison on a small set of records from a pilot that involved CKD-stage-specific cohort selection. Using ICD codes alone, the cohort included a proportion of patients whose most recent GFR values, if used as the staging criterion, would have placed them in a different stage than the code indicated. Adding note extraction to verify the clinician-stated stage and the most recent GFR interpretation reduced that discordance substantially. We are not stating precise numbers from that pilot here because it was a small set and not a published validation study. But the directional result was consistent with what the published phenotyping literature predicts for EHR-based stage assignment.
The Limitation You Still Need to Acknowledge
Note extraction improves cohort definition accuracy for the patients whose notes you can process. It does not solve completeness problems. If a patient received diagnosis-relevant care at a different institution and those records are not in your EHR extract, no extraction pipeline will find those events. The clinical note is only as complete as the EHR record it is drawn from. For the specific task of bridging between code and phenotype within a single health system's data, automated extraction from notes is a genuine improvement. For the broader problem of longitudinal care fragmentation across systems, it is a partial answer at best.
That boundary is worth stating clearly, because overselling what note extraction can do for cohort definition quality leads to studies that address the wrong problems. The most defensible approach is to use extraction to improve what the structured data alone cannot provide, document where you used it and why, validate a sample of your extraction-confirmed cohort against manual review, and acknowledge the completeness ceiling separately as a study limitation. That combination of transparency and tooling is where RWE teams are moving, and the infrastructure to support it is now available at a scale that makes it operationally feasible.