Back

A systematic imputation framework for sparse, multimodal space biology datasets: application to retinal imaging and omics from the RR9 mission

Nagesh, V.; Sanders, L.; Costes, S. V.; Avci, P.; Sigit, A.; Agarwal, A.; Haghighi, A.; Batool, A.; Karouia, F.; Chander, A. M.; Schmidt, C. M.; Gong, J.

2026-06-11 bioinformatics
10.64898/2026.06.09.730965 bioRxiv
Show abstract

Missing data is a fundamental challenge in space biology, where high experimental costs, limited sample availability, and tissue allocation constraints produce datasets that are sparse, multimodal, and heterogeneous. We present a systematic four-stage framework for diagnosing, implementing, and validating data imputation strategies tailored to these characteristics, and demonstrate its application to retinal imaging and omics data from the NASA Rodent Research 9 (RR9) mission. Using logistic regression-based missingness diagnosis, we identify a Missing At Random (MAR) mechanism driven by experimental design constraints across nine assay modalities. We implement and optimize three imputation strategies: K-Nearest Neighbors (KNN), Multiple Imputation by Chained Equations with weak ElasticNet regularization (MICE-Elastic), and a per-column hybrid strategy, evaluated against a random sample imputer baseline. Validation across seven complementary metrics including supervised classification, unsupervised clustering, correlation structure preservation, masked value recovery, cross-dataset generalization, and permutation testing reveals that MICE-Elastic and the Hybrid strategy preserve genuine biological signal in both RNA-seq and TUNEL modalities, while KNN and the random sample imputer do not despite achieving comparable cross-validation accuracy. A critical finding is that imputation substantially improves supervised classification performance while consistently degrading unsupervised clustering structure, a trade-off researchers must understand before applying these methods. This framework provides practical, actionable guidance for space biologists and data scientists managing sparse multimodal datasets, and represents a foundational step toward digital twin development for space medicine.

Matching journals

The top 10 journals account for 50% of the predicted probability mass.

1
PLOS ONE
5266 papers in training set
Top 20%
9.1%
2
Scientific Data
209 papers in training set
Top 0.4%
6.9%
3
GigaScience
212 papers in training set
Top 0.7%
5.0%
4
npj Microgravity
14 papers in training set
Top 0.1%
5.0%
5
Methods in Ecology and Evolution
176 papers in training set
Top 0.5%
5.0%
6
PLOS Computational Biology
1863 papers in training set
Top 8%
4.5%
7
Remote Sensing in Ecology and Conservation
14 papers in training set
Top 0.1%
4.1%
8
Scientific Reports
3612 papers in training set
Top 25%
4.1%
9
Ecology and Evolution
267 papers in training set
Top 2%
3.3%
10
iScience
1154 papers in training set
Top 6%
3.2%
50% of probability mass above
11
Bioinformatics
1204 papers in training set
Top 6%
2.2%
12
PeerJ
308 papers in training set
Top 5%
2.0%
13
Limnology and Oceanography: Methods
11 papers in training set
Top 0.1%
1.5%
14
Briefings in Bioinformatics
354 papers in training set
Top 5%
1.5%
15
Nature Communications
5641 papers in training set
Top 48%
1.4%
16
Communications Biology
993 papers in training set
Top 20%
1.2%
17
Patterns
78 papers in training set
Top 2%
1.2%
18
Frontiers in Bioinformatics
49 papers in training set
Top 0.8%
1.2%
19
Remote Sensing
10 papers in training set
Top 0.1%
1.2%
20
Cell Reports Methods
165 papers in training set
Top 2%
1.2%
21
Journal of Neurotrauma
31 papers in training set
Top 0.4%
1.1%
22
BMC Bioinformatics
457 papers in training set
Top 5%
1.1%
23
Bioinformatics Advances
203 papers in training set
Top 4%
1.1%
24
Genes
144 papers in training set
Top 3%
1.1%
25
Journal of The Royal Society Interface
235 papers in training set
Top 4%
1.0%
26
Microbiology Spectrum
469 papers in training set
Top 9%
1.0%
27
eLife
5828 papers in training set
Top 63%
0.9%
28
Frontiers in Physiology
106 papers in training set
Top 3%
0.9%
29
Microbiome
154 papers in training set
Top 2%
0.9%
30
Genome Biology
637 papers in training set
Top 8%
0.9%