NeuroDiscovery AI database: Comprehensive EHR dataset for Neurology
Selveshwari, S.; Suri, S.; Moodalagiri, S.; Krishna, A.; Kasivajjala, N. C.
Show abstract
PurposeThe NeuroDiscovery AI database is a comprehensive real-world data (RWD) repository containing de-identified electronic health record (EHR) data from U.S.-based neurology outpatient clinics. The structured data encompasses sociodemographic details, clinical examinations, social, medical, and lifestyle histories, International Classification of Diseases (ICD-9/ICD-10) diagnoses, and prescribed medications. Additionally, the database integrates neuroimaging data and laboratory results, providing a robust resource for clinical research. This paper describes a subset of the NeuroDiscovery AI dataset and outlines the processes involved in its development. ParticipantsAs of October 15, 2024, the dataset includes EHR data from 355,791 patients, of whom 40.72% are male. Over 40.06% of the patients are aged 60 or older, spanning across 14,797 distinct diagnosis codes. The data represents more than 15 years of longitudinal patient information, with 26.87% of patients classified as active (defined as having had clinical encounters within the last 18 months). The median follow-up duration for active patients is 19.54 months. LimitationThe large sample size, rigorous data processing, and robust data security of the NeuroDiscovery AI dataset are key strengths, enabling comprehensive studies on disease progression, treatment responses, and long-term outcomes in neurology. The dataset aligns closely with published demographic trends for various neurological conditions, including a female predominance in migraines, multiple sclerosis, and vertigo, with slight variations in age and gender distribution for conditions such as ALS. However, challenges remain, including missing data and data heterogeneity. Ongoing efforts to expand and diversify the dataset aim to improve its applicability and representativeness. Future planThe NeuroDiscovery AI dataset will expand by incorporating data from more providers and improving diversity, aiming to become one of the largest neurology-focused datasets. The platform will continue to evolve into a comprehensive analytical tool, integrating cohort building and data interrogation functionalities to streamline clinical workflows. These enhancements will enable faster, more accurate decision-making, and future efforts will focus on identifying key trends in neurological conditions and patient outcomes.
Matching journals
The top 6 journals account for 50% of the predicted probability mass.
Similar papers in this journal
- A structured course of disease dataset with contact tracing information in Taiwan for COVID-19 modelling 92%
- Construction, Deployment, and Usage of the Human Reference Atlas Knowledge Graph for Linked Open Data 92%
- Anatomical structures, cell types, and biomarkers of the healthy human blood vasculature 92%
Similar papers in this journal
- Large Language Models Facilitate the Generation of Electronic Health Record Phenotyping Algorithms 94%
- PheMIME: An Interactive Web App and Knowledge Base for Phenome-Wide, Multi-Institutional Multimorbidity Analysis 94%
- Increasing Trust in Real-World Evidence Through Evaluation of Observational Data Quality 93%
Similar papers in this journal
- A large dataset of brain imaging linked to health systems data: the curation and access to a whole system national cohort from NHS Scotland 95%
- New implementation of data standards for AI research in precision oncology. Experience from EuCanImage 94%
- An overview of the National COVID-19 Chest Imaging Database: data quality and cohort analysis 93%
Similar papers in this journal
Similar papers in this journal
"Similar papers" are the closest papers from that journal in the model's embedding space. They show what the match is built on, but the ranking comes mostly from a classifier over the whole training set, not from these examples alone.