Back

The WORC* database: MRI and CT scans, segmentations, and clinical labels for 930 patients from six radiomics studies

Starmans, M. P. A.; Timbergen, M. J. M.; Vos, M.; Padmos, G. A.; Grünhagen, D. J.; Verhoef, C.; Sleijfer, S.; van Leenders, G. J. L. H.; Buisman, F. E.; Willemssen, F. E. J. A.; Groot Koerkamp, B.; Angus, L.; van der Veldt, A. A. M.; Rajicic, A.; Odink, A. E.; Renckens, M.; Doukas, M.; de Man, R. A.; IJzermans, J. N. M.; Miclea, R. L.; Vermeulen, P. B.; Thomeer, M. G.; Visser, J. J.; Niessen, W. J.; Klein, S.

2021-08-25 oncology
10.1101/2021.08.19.21262238 medRxiv
Show abstract

The WORC database consists in total of 930 patients composed of six datasets gathered at the Erasmus MC, consisting of patients with: 1) well-differentiated liposarcoma or lipoma (115 patients); 2) desmoid-type fibromatosis or extremity soft-tissue sarcomas (203 patients); 3) primary solid liver tumors, either malignant (hepatocellular carcinoma or intrahepatic cholangiocarcinoma) or benign (hepatocellular adenoma or focal nodular hyperplasia) (186 patients); 4) gastrointestinal stromal tumors (GISTs) and intra-abdominal gastrointestinal tumors radiologically resembling GISTs (246 patients); 5) colorectal liver metastases (77 patients); and 6) lung metastases of metastatic melanoma (103 patients). For each patient, either a magnetic resonance imaging (MRI) or computed tomography (CT) scan, collected from routine clinical care, one or multiple (semi-)automatic lesion segmentations, and ground truth labels from a gold standard (e.g., pathologically proven) are available. All datasets are multicenter imaging datasets, as patients referred to our institute often received imaging at their referring hospital. The dataset can be used to validate or develop radiomics methods, i.e., using machine or deep learning to relate the visual appearance to the ground truth labels, and automatic segmentation methods. See also the research article related to this dataset: Starmans et al., Reproducible radiomics through automated machine learning validated on twelve clinical applications, Submitted. Specifications Table O_TBL View this table: org.highwire.dtl.DTLVardef@b12793org.highwire.dtl.DTLVardef@9d440corg.highwire.dtl.DTLVardef@dea161org.highwire.dtl.DTLVardef@352338org.highwire.dtl.DTLVardef@9b5483_HPS_FORMAT_FIGEXP M_TBL C_TBL

Matching journals

The top 6 journals account for 50% of the predicted probability mass.

50% of probability mass above

"Similar papers" are the closest papers from that journal in the model's embedding space. They show what the match is built on, but the ranking comes mostly from a classifier over the whole training set, not from these examples alone.