Back

CAT-APP: Contamination Analysis and Tempering--An Automated Online Platform for Plasma Proteomics with Data Rescuing Capabilities

Niu, M.; Zhang, D.; Zhou, Z.; Zhang, J.; Zhang, R.; Shen, L.; Wang, H.

2025-07-11 bioinformatics
10.1101/2025.07.08.663798 bioRxiv
Show abstract

Plasma and serum proteome profiling have become central to biomarker discovery for medical research. The advent of high-throughput plasma proteomics pipelines has enabled the generation of massive datasets. However, biomarker discovery is frequently compromised by sample contamination, primarily from erythrocytes, platelets, and coagulation factors. Such contamination can lead to the false identification of contaminants as biomarkers in over 50% of studies or obscure genuine biomarkers due to contamination-induced variability. Currently, no tools are available to salvage such compromised data, and the default approach is often to discard the affected samples. To address this gap, we introduce CAT-APP (Contamination Analysis and Tempering-An Automated Online Platform for Plasma Proteomics). CAT-APP tackles contamination through three key modules: multi-dimensional contamination assessment and adaptive contamination indexing, mathematical model-based contamination tempering, and data recovery evaluation with visualization. We applied CAT-APP to 28 independent plasma proteomics datasets and found that 82% exhibited contamination issues. CAT-APP markedly improved data quality by removing false-positive biomarkers introduced by contamination and restoring true biological signals that were previously obscured in representative datasets. To promote widespread adoption, we provide CAT-APP as an freely accessible, user-friendly, and robust web platform at https://www.bloodecosystem.com/tools/CAT-APP/.

Matching journals

The top 3 journals account for 50% of the predicted probability mass.

50% of probability mass above

"Similar papers" are the closest papers from that journal in the model's embedding space. They show what the match is built on, but the ranking comes mostly from a classifier over the whole training set, not from these examples alone.