Back

Natural language processing and modeling of clinical disease trajectories across brain disorders

Mekkes, N. J.; Groot, M.; Wehrens, S.; Hoekstra, E. J.; Herbert, M.; Brummer, M.; Wever, D.; Netherlands Neurogenetics Database consortium, ; Rozenmuller, A.; Huitinga, I.; Holtman, I. R.

2022-09-23 health informatics
10.1101/2022.09.22.22280158 medRxiv
Show abstract

Brain disorders, including neurodegenerative diseases and mental illnesses, are often difficult to diagnose and study due to clinical and pathological heterogeneity, overlap in clinical manifestations between disorders, and frequent comorbidities, hampering drug development and fundamental research. Hence, there is a clear need for data-driven approaches to disentangle these complex disorders. Here, we established a computational pipeline to process clinical summaries from donors with a wide range of brain disorders that were neuropathologically diagnosed by the Netherlands Brain Bank. First, we identified and defined 90 cross-disorder signs and symptoms within cognitive, motor, sensory, psychiatric, and general domains. Second, we trained and optimized natural language processing (NLP) models to identify these signs and symptoms in individual sentences of the extensive clinical summaries from donors of the NBB, resulting in temporal disease trajectories. Third, we studied the temporal manifestation and survival profiles across rare and complex dementias, alpha-synucleinopathies, frontotemporal dementia subtypes, and mental illnesses, giving new insight into how symptomatology differs in manifestation and temporal profiles across brain disorders. Lastly, we trained a recurrent neural network to predict the Neuropathological Diagnosis. Taken together, this integrated approach resulted in a highly unique resource that can facilitate research into cross-disorder symptomatology.

Matching journals

The top 6 journals account for 50% of the predicted probability mass.

50% of probability mass above

"Similar papers" are the closest papers from that journal in the model's embedding space. They show what the match is built on, but the ranking comes mostly from a classifier over the whole training set, not from these examples alone.