Tracking patient clusters over time enables to extract all the information available in the medico-administrative databases
LAMBERT, J.; LEUTENEGGER, A.-L.; JANNOT, A.-S.; BAUDOT, A.
Show abstract
ContextIdentifying clusters (i.e., subgroups) of patients from the analysis of medico-administrative databases is particularly important to better understand disease heterogeneity. However, the complexity of these databases, in particular due to the presence of truncated longitudinal data, requires adaptation of clustering approaches. ObjectiveWe propose here cluster-tracking approaches to identify clusters of patients from longitudinal data contained in medico-administrative databases. Material and MethodsWe first cluster patients at each age using either the Markov Cluster algorithm (MCL) from patient networks or Kmeans from raw data. We then track the identified clusters over ages to construct cluster-trajectories. We compared our novel approaches with three longitudinal clustering approaches by calculating the silhouette score. As a use-case, we analyzed antithrombotic drugs prescribed from 2008 to 2018 contained in the Echantillon Generaliste des Beneficiaires (EGB), a French national cohort. ResultsOur cluster-tracking approaches allowed us to identify several cluster-trajectories having clinical significance. Silhouette score comparison between the different approaches reveals that the best score is obtained for the cluster-tracking approaches. ConclusionThe cluster-tracking approaches are a novel and efficient alternative to identify patient clusters from medico-administrative databases by taking into account their specificities.
Matching journals
The top 6 journals account for 50% of the predicted probability mass.
Similar papers in this journal
Similar papers in this journal
Similar papers in this journal
- Deep representation learning for clustering longitudinal survival data from electronic health records 94%
- Projecting genetic associations through gene expression patterns highlights disease etiology and drug mechanisms 93%
- CellScope: High-Performance Cell Atlas Workflow with Tree-Structured Representation 93%
Similar papers in this journal
- An algorithm to build synthetic temporal contact networks based on close-proximity interactions data 94%
- Multi-omics subtyping of hepatocellular carcinoma patients using a Bayesian network mixture model 93%
- Biological networks and GWAS: comparing and combining network methods to understand the genetics of familial breast cancer susceptibility in the GENESIS study 92%
"Similar papers" are the closest papers from that journal in the model's embedding space. They show what the match is built on, but the ranking comes mostly from a classifier over the whole training set, not from these examples alone.