Back

From free text to SOFA score: automated reconstruction of sepsis severity from unstructured clinical notes

Monsalve Barrientos, K.; Castano-Villegas, N.; Escalon, E. R.; Zea, J.; Velasquez, L.

2025-12-19 intensive care and critical care medicine
10.64898/2025.12.17.25342509 medRxiv
Show abstract

ObjectiveTo evaluate the ability of a natural language processing system to automatically reconstruct the SOFA score from unstructured clinical notes in patients with sepsis and validate its applicability in intensive care units. Materials and methodsRetrospective study in the MIMIC-III database that included 284 adults with sepsis. The SOFA calculated with structured data was compared with the SOFA reconstructed by free text extraction. Clinical rules were applied for calculation at 24 h and 48 h. Variable completeness, severity reclassification, and association with hospital mortality were evaluated using logistic regression. ResultsAutomated extraction increased the availability of critical variables (respiratory 33% to 100%, vasopressor 12% to 41%). The reconstructed SOFA increased by 3 points at 24 hours, reclassifying patients with high severity (SOFA [&ge;] 6) from 17% to 48% and SOFA [&ge;] 10 from 5% to 22%. Reconstructed scores remained associated with mortality at 24 h (OR 1.16, 95% CI 1.09-1.24) and at 48 h (OR 1.23, 95% CI 1.15-1.31), comparable to that based on structured data (p < 0.001). DiscussionAutomatic reconstruction of the SOFA from free text recovers information missing from structured fields, reducing underestimation of severity. ConclusionNLP approaches supported by large language models provide a more complete and clinically consistent SOFA score in sepsis when structured data are insufficient.

Matching journals

The top 7 journals account for 50% of the predicted probability mass.

50% of probability mass above

"Similar papers" are the closest papers from that journal in the model's embedding space. They show what the match is built on, but the ranking comes mostly from a classifier over the whole training set, not from these examples alone.