Clinical utility of LLM-assisted chart review for the detection of bleeding events - 30/08/26
, Oscar M. van der Meer c, Alberto Testoni d, Pieter O.L.P. Broeren a, Dave A. Dongelmans a, Paul W.G. Elbers a, Marcella C.A. Müller a, Senta Jorinde Raasveld a, b, Alexander P.J. Vlaar aHighlights |
• | Bleeding assessment tools often require manual chart review. |
• | Large Language Models (LLMs) may help identify bleeding events. |
• | The LLM detected more bleeding events than prior manual review. |
• | The LLM frequently misclassified bleeding site and severity. |
• | Expert adjudication remains necessary for accurate classification. |
Abstract |
Introduction |
Red blood cell (RBC) transfusions are frequently administered in the intensive care unit (ICU) and are independently associated with increased mortality. While most studies have focused on transfusion events, bleeding events that do not result in transfusion remain understudied. The validated HEME scoring system enables structured bleeding assessment but requires labor-intensive chart review, limiting scalability. We hypothesized that large language models (LLMs) could function as screening instruments to facilitate scalable chart review and support the development of curated bleeding event datasets.
Methods |
This retrospective cohort study included critically ill patients at high risk of bleeding. Reference data on in-ICU bleeding events were derived from prior manual review using the Hemorrhage Measurement (HEME) scoring system. For the LLM-based analysis, unstructured clinical notes were extracted from the electronic health record (EHR) for the same patient cohort. Notes were processed via structured JSON-formatted prompts using the OpenAI gpt-4o-mini model in a secure environment. Next, these bleeding events were adjudicated to come to a verified reference set.
Results |
In 149 patients 36 490 notes were analyzed. A total of 654 adjudicated true bleeding events were identified in the reference set. Manual review detected 66 events, while the LLM detected 647 events. Manual review identified 7 true events missed by the LLM 1.1% (7/654). The LLM identified 588 true events not captured by manual review, corresponding to an incremental detection yield of 90.1% (588/654). Among the 739 events identified by the LLM, 85 were not confirmed upon adjudication, resulting in a false-positive rate of 11.5% (85/739). The LLM correctly classified bleeding site in 42% of cases (272/654) and bleeding severity in 65% of cases (424/654).
Conclusion |
Use of an LLM substantially increased detection of ICU bleeding events compared with manual review. However, an 11.5% false-positive rate and limited accuracy in site and severity classification indicate that expert adjudication remains necessary.
Le texte complet de cet article est disponible en PDF.Keywords : Transfusion medicine, Large language models, Bleeding assessment, Artificial intelligence, Intensive care
Plan
Bienvenue sur EM-consulte, la référence des professionnels de santé.
L’accès au texte intégral de cet article nécessite un abonnement.
Déjà abonné à cette revue ?
