Accessible and Reproducible Renal Cell Carcinoma Research Through Open-Sourcing Data and Annotations
de Boer, S.; Häntze, H.; Ziegelmayer, S.; van Ginneken, B.; Prokop, M.; Bressem, K. K.; Hering, A.
Show abstract
BackgroundMedical imaging, especially computed tomography and magnetic resonance imaging, is essential in clinical care of patients with renal cell carcinoma (RCC). Artificial intelligence (AI) research into computer-aided diagnosis, staging and treatment planning needs curated and annotated datasets. Across literature, The Cancer Genome Atlas (TCGA) datasets are widely used for model training and validation. However, re-annotation is often necessary due to limited access to public annotations, raising entry barriers and hindering comparison with prior work. MethodsWe screened 1915 CT scans from three TCGA-RCC databases and employed a segmentation model to annotate kidney lesion. After a meta-data-based exclusion step, we hosted a reader study with all papillary (n=56), chromophobe (n=27) and 200 randomly selected clear cell RCC cases. Two students quality checked and corrected the data as well as annotated tumors and cysts. Uncertain cases were checked by a board-certified radiologist. ResultsAfter data exclusion and quality control a total of 142 annotated CT scans from 101 patients (26 female, 75 male, mean age 56 years) remained. This includes 95 CTs with clear cell RCC, 29 with papillary RCC and 18 with chromophobe RCC. Images and voxel-level annotations of kidneys and lesions are open sourced at https://zenodo.org/records/19630298. ConclusionBy making the annotations open-source, we encourage accessible and reproducible AI research for renal cell carcinoma. We invite other researchers who have previously annotated any of these cohorts to share their annotations.
Matching journals
The top 6 journals account for 50% of the predicted probability mass.
Similar papers in this journal
- ai-corona : Radiologist-Assistant Deep Learning Framework for COVID-19 Diagnosis in Chest CT Scans 93%
- Classification performance bias between training and test sets in a limited mammography dataset 93%
- Enhancing Semantic Segmentation in Chest X-Ray Images through Image Preprocessing: ps-KDE for Pixel-wise Substitution by Kernel Density Estimation 93%
Similar papers in this journal
- A deep learning approach for Pan-Renal Cell Carcinoma classification and survival prediction from histopathology images 93%
- Segmentation of Pancreatic Ductal Adenocarcinoma (PDAC) and surrounding vessels in CT images using deep convolutional neural networks and Texture Descriptors 93%
- CluSA: Clustering-based Spatial Analysis framework through Graph Neural Network for Chronic Kidney Disease Prediction using Histopathology Images 93%
Similar papers in this journal
- Assessing GPT-4 Multimodal Performance in Radiological Image Analysis 94%
- Evaluating Large Language Model-Generated Brain MRI Protocols: Performance of GPT4o, o3-mini, DeepSeek-R1 and Qwen2.5-72B 92%
- From Community Acquired Pneumonia to COVID-19: A Deep Learning Based Method for Quantitative Analysis of COVID-19 on thick-section CT Scans 91%
Similar papers in this journal
- Deep Learning Segmentation of Glomeruli on Kidney Donor Frozen Sections 93%
- “E Pluribus Unum”: Prospective acceptability benchmarking from the Contouring Collaborative for Consensus in Radiation Oncology (C3RO) Crowdsourced Initiative for Multi-Observer Segmentation 92%
- Predicting Primary Site of Secondary Liver Cancer with a Neural Estimator of Metastatic Origin (NEMO) 91%
Similar papers in this journal
- Fully Automated Explainable Abdominal CT Contrast Media Phase Classification Using Organ Segmentation and Machine Learning 95%
- Phase Recognition in Contrast-Enhanced CT Scans based on Deep Learning and Random Sampling 93%
- PSMA-Hornet: fully-automated, multi-target segmentation of healthy organs in PSMA PET/CT images 91%
"Similar papers" are the closest papers from that journal in the model's embedding space. They show what the match is built on, but the ranking comes mostly from a classifier over the whole training set, not from these examples alone.