Back

A central research portal for mining pancreatic clinical and molecular datasets and accessing biobanked samples

Oscanoa, J.; Ross-Adams, H.; Dayem Ullah, A. Z. M.; Kolvekar, T. S.; Sivapalan, L.; Gadaleta, E.; Thorn, G. J.; Abdollahyan, M.; Imrali, A.; Saad, A.; Roberts, R.; Hughes, C.; PCRFTB, ; Kocher, H. M.; Chelala, C.

2024-07-26 oncology
10.1101/2024.07.25.24309825 medRxiv
Show abstract

The Pancreas Genome Phenome Atlas (PGPA) is dedicated to the analysis of pancreatic datasets from four primary sources (Cancer Genome Atlas, International Cancer Genome Consortium, Cancer Cell Line Encyclopaedia, Genomics Evidence Neoplasia Information Exchange) that together form the foundation of -omics profiling of pancreatic malignancies and related lesions (n=7,760 specimens). Multiple user-friendly analytical tools to explore the associated molecular data from these primary specimens and cell lines are available. Crucially, PGPA is the access point for Pancreatic Cancer Research Fund Tissue Bank - the only national pancreatic cancer biobank in the UK, and will facilitate effective sharing of multi-modal molecular, histopathology and imaging data from biobank samples (>60,000 specimens from >3,400 cases and controls; 2,037 H&E images from 349 donors) and accelerate validation of in silico findings in patient-derived material. This places PGPA at the forefront of biomarker-based research, providing the user community with a distinct resource to facilitate hypothesis-testing on public data, validate novel research findings, and access curated, high-quality patient tissues for translational research. To demonstrate the practical utility of PGPA, we investigate somatic variants associated with established transcriptomic subtypes and disease prognosis: several patient-specific variants are clinically actionable and may be leveraged for precision medicine.

Published in Translational Oncology (predicted rank #10) · training set

Matching journals

The top 7 journals account for 50% of the predicted probability mass.

50% of probability mass above

"Similar papers" are the closest papers from that journal in the model's embedding space. They show what the match is built on, but the ranking comes mostly from a classifier over the whole training set, not from these examples alone.