Back

Contextual Evaluation of MicroRNA Sequencing Data Harmonization: Performance in Sample Clustering

Zou, J.; Duren, Y.; Wang, X.; Xiang, Y.; Qi, Y.; Wang, M.; Wu, Y.; Singer, S.; Qin, L.-X.

2026-08-14 bioinformatics
10.64898/2026.08.09.743718 bioRxiv
Show abstract

Reliable translation of microRNA sequencing data depends on effective harmonization to mitigate artifacts from variable experimental handling. Although many harmonization methods exist, prior evaluations have focused mainly on differential expression analysis, leaving the impact on subgroup discovery understudied. We present a framework for evaluating harmonization in the context of sample clustering, which integrates AI-augmented datasets, statistical evaluation pipelines, and accessible software tools, enabling systematic comparisons across diverse signal-to-artifact ratios and cluster composition settings. Using this framework, we show that harmonization can, often partially, restore clustering accuracy lost to artifacts, especially at moderate signal-to-artifact ratios, with the level of gains depending on the specific harmonization method, the paired clustering technique, and the cluster composition setting. We further confirm these findings by analyzing reconstructed cohorts from The Cancer Genome Atlas breast cancer microRNA sequencing data. Collectively, the results underscore the need for tailored harmonization to support reliable subgroup discovery and highlight the broader importance of context-specific workflows in translational genomics.

Matching journals

The top 5 journals account for 50% of the predicted probability mass.

50% of probability mass above

"Similar papers" are the closest papers from that journal in the model's embedding space. They show what the match is built on, but the ranking comes mostly from a classifier over the whole training set, not from these examples alone.