Quantifying data reuse in proteomics using PRIDE downloads statistics and a semi-supervised LLM-based framework
Hewapathirana, S.; Bai, J.; Bandla, C.; Kamatchinathan, S.; Kundu, D. J.; John, N. S.; Brown-Harry, B.; Madhusoodanan, N.; Riera Duocastella, J. M.; Vizcaino, J. A.; Perez-Riverol, Y.
Show abstract
Understanding how scientific datasets are accessed and reused is essential for resource planning and impact assessment. Here we present the PRIDE Archive download tracking infrastructure and a comprehensive analysis of 159.3 million download records from the PRIDE proteomics database (2021-2025), spanning 35,528 datasets accessed from 235 locations. The infrastructure includes nf-downloadstats, a scalable Nextflow pipeline for processing download logs, and DeepLogBot, a machine-learning framework that classifies traffic into bots, institutional download hubs, and independent user downloads. DeepLogBot combines heuristic seed selection with multi-LLM annotation (Claude and Qwen3) to produce gold-standard training labels, achieving 92.2% bot classification accuracy on a held-out test set. After separating bot traffic, analysis reveals downloads from 214 countries/regions, 249 institutional download hubs, and a concentrated reuse distribution, with the top five countries (United States, United Kingdom, Germany, China, and Canada) accounting for over 54% of independent user downloads. These findings provide actionable insights for repository infrastructure planning and highlight the importance of distinguishing automated from individual access in scientific data resources.
Matching journals
The top 3 journals account for 50% of the predicted probability mass.
Similar papers in this journal
- An adaptive, continuous-learning framework for clinical decision-making from proteome-wide biofluid data 94%
- LEOPARD: missing view completion for multi-timepoints omics data via representation disentanglement and temporal knowledge transfer 94%
- Imputation of label-free quantitative mass spectrometry-based proteomics data using self-supervised deep learning 94%
Similar papers in this journal
- PEPerMINT: Peptide Abundance Imputation in Mass Spectrometry-based Proteomics using Graph Neural Networks 94%
- SHEPHARD: a modular and extensible software architecture for analyzing and annotating large protein datasets 93%
- FAVA: High-quality functional association networks inferred from scRNA-seq and proteomics data 93%
Similar papers in this journal
"Similar papers" are the closest papers from that journal in the model's embedding space. They show what the match is built on, but the ranking comes mostly from a classifier over the whole training set, not from these examples alone.