Back

The usage of transcriptomics datasets as sources of Real-World Data for clinical trialling

Matos-Filipe, P.; Garcia-Illarramendi, J. M.; Jorba, G.; Oliva, B.; Farres, J.; Mas, J. M.

2022-11-13 bioinformatics
10.1101/2022.11.10.515995 bioRxiv
Show abstract

BackgroundRandomised Clinical Trials (RCT) reflect results within their specific controlled settings, necessitating further studies to understand outcomes across all possible scenarios. The usage of Real-World Data (RWD) has been recently considered to be a viable alternative to overcome these issues and complement clinical conclusions. Molecular profiles of patients captured by high-throughput measures reflect their medical conditions. When this information is linked to clinical and demographical information, nuances in transcriptomics data can uncover subtle variations in disease pathways among distinct patient groups. This work focuses on the construction of a patient repository database with molecular and clinical information resulting from the integration of publicly available transcriptomics datasets. ResultsPatient data were integrated into the patient repository by using a novel post-processing technique allowing for the usage of samples originating from different/multiple Gene Expression Omnibus (GEO) datasets. Our post-processing technique, which we have named MicroArray Cross-plAtfoRm pOst-prOcessiNg (MACAROON), aims to standardise and integrate transcriptomics data (considering batch effects and possible processing-originated artefacts). This process was able to better reproduce the down streaming biological conclusions in a 45% improvement compared to other methods available. Furthermore, RWD was mined from GEO samples metadata and a clinical and demographical characterisation of the database was obtained. RWD mining was done through a manually curated synonym dictionary allowing for the correct assignment (95.33% median accuracy; 80.14% average) of medical conditions. ConclusionsOur strategy produced a repository, which includes molecular, clinical and demographical RWD by integrating multiple public datasets. The exploration of these data facilitates the discovery of clinical outcomes and molecular pathways specific to predetermined patient populations.

Matching journals

The top 6 journals account for 50% of the predicted probability mass.

50% of probability mass above

"Similar papers" are the closest papers from that journal in the model's embedding space. They show what the match is built on, but the ranking comes mostly from a classifier over the whole training set, not from these examples alone.