Using large language models to enhance clinically-driven missing data recovery algorithms in electronic health records

Objectives Electronic health record (EHR) data are prone to missingness and errors. Previously, we devised an enriched chart review protocol where a “roadmap” of auxiliary diagnoses was used to recover missing values. Still, chart reviews are expensive and time-intensive, limiting the number of patients whose data can be reviewed. Now, we investigate the accuracy and scalability of a roadmap-driven algorithm, based on International Classification of Diseases, 10th revision (ICD-10) codes, to mimic expert chart reviews and recover missing values. Materials and Methods In addition to the clinicians’ original roadmap from our previous work, we consider new versions that were iteratively refined using large language models (LLMs) in conjunction with clinical expertise to expand the list of auxiliary diagnoses. Using chart reviews for 100 patients from an extensive EHR, we examine algorithm performance. Results Across 100 chart reviewed patients, there were 413 missing values in the EHR data. The expert chart reviews recovered 49 (12%), while the algorithms using LLMs-enhanced roadmaps recommended almost twice as many (83-89, 20%-22%). The final algorithm using clinician-approved LLMs’ additions offered a balance (73, 18%), expanding the original roadmap with LLMs’ suggestions but only when deemed clinically relevant. When applied to a larger study of 1000 patients from the same EHR, the per-patient median number of non-missing values increased from 6 to 7. Discussion Clinically-driven algorithms (enhanced by LLMs) can recover missing EHR data with similar accuracy to chart reviews and feasibly be applied to large samples. Extending them to monitor other data quality dimensions is a promising future direction.

Posted on:
June 2, 2026
Length:
2 minute read, 254 words
Categories:
Peer Reviewed Article
See Also: