GROMTools: scalable individual-level GReX imputation for mega-biobank-scale cohorts
Anyfantakis, M.; Venkatesh, S.; Bennett, J. J. R.; Hoffman, G. E.; Roussos, P.; Voloudakis, G.
Show abstract
MotivationThe computational burden of individual-level genetically regulated gene expression (GReX) imputation has risen sharply with the growth of human mega-biobanks and the rapid expansion of transcriptomic imputation models across tissues and single-cell hierarchies. Existing tools were not designed for this setting and require complex, memory-intensive workflows that are poorly matched to shared and cloud-based compute environments, where runtime, memory, and I/O directly determine cost and throughput. GROMTools is an open-source C++ engine with an R interface that exploits sparse prediction weights, streams PLINK2 genotypes, and writes compact binary outputs for scalable individual-level GReX imputation. ResultsIn UK Biobank (UKBB), benchmarks on chromosome 1 across 50,000-450,000 individuals, 388,017 variants, and 11,724 gene-tissue pairs from 32 single-cell models, GROMTools produced near-identical predictions to PrediXcan and PLINK2, with minimum Pearson correlation >0.999 and maximum RMSE <0.001 across all of the imputed genes, while reducing CPU time by about 100-fold and peak memory by about 33-fold. These gains make routine biobank-scale individual-level GReX imputation practical and cost-efficient on standard CPU infrastructure. AvailabilityGROMTools is freely available at https://github.com/voloudakislab/gromtools under the GPL v3 license. Documentation available at https://voloudakislab.github.io/gromtools/. All of our coding scripts used for quality control of data and benchmarking pipelines and all of our log files are archived at DOI: https://doi.org/10.5281/zenodo.19547333.
Matching journals
The top 1 journal accounts for 50% of the predicted probability mass.
Similar papers in this journal
Similar papers in this journal
"Similar papers" are the closest papers from that journal in the model's embedding space. They show what the match is built on, but the ranking comes mostly from a classifier over the whole training set, not from these examples alone.