The usage of transcriptomics datasets as sources of Real-World Data for clinical trialling
Matos-Filipe, P.; Garcia-Illarramendi, J. M.; Jorba, G.; Oliva, B.; Farres, J.; Mas, J. M.
Show abstract
BackgroundRandomised Clinical Trials (RCT) reflect results within their specific controlled settings, necessitating further studies to understand outcomes across all possible scenarios. The usage of Real-World Data (RWD) has been recently considered to be a viable alternative to overcome these issues and complement clinical conclusions. Molecular profiles of patients captured by high-throughput measures reflect their medical conditions. When this information is linked to clinical and demographical information, nuances in transcriptomics data can uncover subtle variations in disease pathways among distinct patient groups. This work focuses on the construction of a patient repository database with molecular and clinical information resulting from the integration of publicly available transcriptomics datasets. ResultsPatient data were integrated into the patient repository by using a novel post-processing technique allowing for the usage of samples originating from different/multiple Gene Expression Omnibus (GEO) datasets. Our post-processing technique, which we have named MicroArray Cross-plAtfoRm pOst-prOcessiNg (MACAROON), aims to standardise and integrate transcriptomics data (considering batch effects and possible processing-originated artefacts). This process was able to better reproduce the down streaming biological conclusions in a 45% improvement compared to other methods available. Furthermore, RWD was mined from GEO samples metadata and a clinical and demographical characterisation of the database was obtained. RWD mining was done through a manually curated synonym dictionary allowing for the correct assignment (95.33% median accuracy; 80.14% average) of medical conditions. ConclusionsOur strategy produced a repository, which includes molecular, clinical and demographical RWD by integrating multiple public datasets. The exploration of these data facilitates the discovery of clinical outcomes and molecular pathways specific to predetermined patient populations.
Matching journals
The top 6 journals account for 50% of the predicted probability mass.
Similar papers in this journal
Similar papers in this journal
Similar papers in this journal
- Topological embedding and directional feature importance in ensemble classifiers for multi-class classification 93%
- Gra-CRC-miRTar: The pre-trained nucleotide-to-graph neural networks to identify potential miRNA targets in colorectal cancer 93%
- DeepCORE: An interpretable multi-view deep neural network model to detect co-operative regulatory elements 93%
Similar papers in this journal
- NetActivity enhances transcriptional signals by combining gene expression into robust gene set activity scores through interpretable autoencoders 96%
- OmicsFootPrint: a framework to integrate and interpret multi-omics data using circular images and deep neural networks 95%
- Assessing the impact of transcriptomics data analysis pipelines on downstream functional enrichment results 94%
"Similar papers" are the closest papers from that journal in the model's embedding space. They show what the match is built on, but the ranking comes mostly from a classifier over the whole training set, not from these examples alone.