Back

Improving inferences regarding individual patients using Emergency Medical Services response-based data in the National Emergency Medical Services Information System (NEMSIS)

Morrison, C. N.; Bushover, B. R.; Crowe, R. P.; Mills, C. W.; Lo, A. X.; Rundle, A. G.

2025-09-05 emergency medicine
10.1101/2025.09.02.25334950 medRxiv
Show abstract

Study ObjectiveThe National Emergency Medicine Services Information System (NEMSIS) public release dataset is an important tool for researching Emergency Medical Services (EMS) responses. However, multiple EMS units might attend to a single patient and the data are organized by EMS response, presenting challenges to inferences about patient-level events. We test whether data on time of the 911 call and patient characteristics can be used to screen for multiple EMS records that reflect a single patient encounter. MethodsUsing data on EMS responses to assaults in New York City in 2024 we identified EMS responses that had identical data for time of 911 call, patient age, sex, race/ethnicity, and longitude-latitude where the patient was encountered. EMS responses with identical data for all of these variables were assumed to have attended to a single patient. We then assessed the validity of using matches on 911 call time, patient age, sex, race/ethnicity (i.e., without latitude-longitude) to identify instances where separate EMS responses matched for these variables plus location. ResultsOf 32,202 EMS responses, 5,143 responses matched other responses for all variables, suggesting that there were 26,451 patients encounters. Matching on permutations of variables for time of 911 call, patient age, sex, race and ethnicity had 100% sensitivity and a high specificity (range 91.3% to 98.6%) for identifying responses that matched on all of these variables plus longitude-latitude. ConclusionData available in the NEMSIS public release dataset can be used to screen for duplicate EMS responses improving inferences about patient level events.

Matching journals

The top 3 journals account for 50% of the predicted probability mass.

50% of probability mass above

"Similar papers" are the closest papers from that journal in the model's embedding space. They show what the match is built on, but the ranking comes mostly from a classifier over the whole training set, not from these examples alone.