MPGEM: A harmonized and transcriptome-complete resource for large-scale reuse of legacy human microarray data
Gupta, S.; Verma, A. K.; Jana, S.; Ahmad, S.
Show abstract
Abstract Background: Legacy microarray datasets provide an extensive record of human transcriptomic biology, but their reuse is constrained by differences in platform design, preprocessing, measurement scale, and gene coverage. Platforms measuring only subsets of genes cannot readily be integrated with higher-coverage platforms, limiting large-scale analysis and computational modeling. Results: We developed Multi-Platform Gene Expression Matrix (MPGEM), a computational framework and resource for harmonizing and completing gene-expression profiles across heterogeneous microarray platforms. MPGEM uses a Reference Quantile Distribution (RQD) and generalized Reference Subset Quantile Distribution (RSQD) framework to transform profiles with different gene coverage onto a common quantitative scale. The MPGEM Engine, a multilayer perceptron, predicts expression of unmeasured genes from genes shared across platforms. Applied to Affymetrix GPL570, GPL571, and GPL96, MPGEM uses GPL570 as a 19,320- gene reference space comprising 12,712 predictor and 6,608 target genes. The resulting resource contains 207,135 human gene-expression profiles across 19,320 genes. Evaluation using masked GPL570 profiles yielded mean sample-wise Pearson and Spearman correlations of 0.944 and 0.939, respectively, and mean gene-wise correlations of 0.830 and 0.825. The lowest-performing 5% of target genes achieved a mean Pearson correlation of 0.683. MPGEM showed comparable or higher predictive performance than baseline mean imputation and K-nearest-neighbor approaches. Conclusions: MPGEM transforms heterogeneous, partially measured legacy microarray profiles into a harmonized, transcriptome-complete representation, facilitating their reuse for large-scale transcriptomic analysis, biomarker discovery, systems biology, and machine learning. The framework, trained models, and expression resource are provided as open-source resources.
Matching journals
The top 4 journals account for 50% of the predicted probability mass.
Similar papers in this journal
- Extraction of biological terms using large language models enhances the usability of metadata in the BioSample database 94%
- GEfetch2R: fetching single-cell/bulk RNA-seq data from public repositories to R and benchmarking the subsequent format conversion tools 93%
- TuBA: Tunable Biclustering Algorithm Reveals Clinically Relevant Tumor Transcriptional Profiles in Breast Cancer 92%
Similar papers in this journal
- SummArIzeR: Simplifying cross-database enrichment result clustering and annotation via large language models 94%
- CuBlock: A cross-platform normalization method for gene-expression microarrays 94%
- ExpOmics: a comprehensive web platform empowering biologists with robust multi-omics data analysis capabilities 93%
Similar papers in this journal
- Differential Expression Analysis with InMoose, the Integrated Multi-Omic Open-Source Environment in Python 93%
- GEOlimma: Differential Expression Analysis and Feature Selection Using Pre-Existing Microarray Data 93%
- DAESC+: High-performance, integrated software for single-cell allele-specific expression data 93%
Similar papers in this journal
"Similar papers" are the closest papers from that journal in the model's embedding space. They show what the match is built on, but the ranking comes mostly from a classifier over the whole training set, not from these examples alone.