Research on Crop Phenotype Prediction Methods Based on SNP-context and Whole-genome Features Embedding
Huan, L.; peng, C. Y.; Zhen, C.; Ting, W.; Chao, W.; bo, B. W.; Juan, L.; Mo, W.; Li, C.; ming, W. J.
Show abstract
Modern agriculture demands precise genomic prediction to accelerate elite crop breeding, yet traditional genomic prediction approaches, such as genomic best linear unbiased prediction (GBLUP) and Bayesian methods, focus primarily on the cumulative effect of individual SNPs, thus neglecting the concerted influence that the surrounding sequence context has on the phenotype. To overcome these limitations, we propose two novel feature embedding modes (SNP-context and whole-genome) based on DNABERT-2, a cross-species genomic foundation model that uses self-attention mechanisms and transfer learning to automatically identify conserved sequence features across diverse evolutionary lineages without prior biological assumptions. The whole-genome feature embedding aggregates genomic information at a global scale by pooling vectors from chunked sequences processed by DNABERT-2, whereas the context feature embedding captures local information by directly encoding variable-length (500--3000 bp) sequences centered on target SNPs. To reduce noise in the high-dimensional feature embeddings, we employed principal component analysis (PCA) and partial least squares (PLS) to project the features into a lower-dimensional space. We generated two kinds of feature embedding for three crop datasets (rice413, rice395, and maize301), investigated the impact of 500--3000 bp flanking SNP contexts on phenotypic prediction, and compared prediction accuracy variations across algorithms at 4--768 feature dimensions among the PCA, PLS, and no dimensionality reduction strategies. The results demonstrate that machine learning (ML) algorithms operating under the SNP-context embedding mode achieve greater accuracy and lower mean absolute errors (MAEs) than traditional SNP features do at specific context lengths, particularly for traits with low-to-moderate heritability (h2[isin](0.2, 0.7]). In contrast, using whole-genome embeddings as input for ML can further improve the prediction accuracy for highly heritable traits (h2[isin](0.7, 1.0]), even outperforming state-of-the-art deep learning models (such as DNNGP and ResGS) that rely on SNP markers. Our code is available on https://github.com/oliveSpring/Crop_DNA_Embedding.git
Matching journals
The top 4 journals account for 50% of the predicted probability mass.
Similar papers in this journal
- Genomic and Phenomic Prediction for Soybean Seed Yield, Protein, and Oil 95%
- Multi-Trait Machine and Deep Learning Models for Genomic Selection using Spectral Information in a Wheat Breeding Program 95%
- Leveraging genomics and temporal high-throughput phenotyping to enhance association mapping and yield prediction in sesame 95%
Similar papers in this journal
- Accounting for epistasis improves genomic prediction of phenotypes with univariate and bivariate models across environments 96%
- Comparative analysis of genomic prediction approaches for multiple time-resolved traits in maize 95%
- Bayesian optimization of multivariate genomic prediction models based on secondary traits for improved accuracy gains and phenotyping costs 94%
Similar papers in this journal
- Enviromic assembly increases accuracy and reduces costs of the genomic prediction for yield plasticity 95%
- Incorporating gene expression and environment improves genomic prediction of wheat traits 95%
- Leveraging Transcriptomics-Based Approaches to Enhance Genomic Prediction: Integrating SNP weights and gene-networks for Cotton Fibre Quality Improvement 95%
Similar papers in this journal
Similar papers in this journal
- Characterizing yield through wheat's perception of chronological progression: a multi-omics plant-time warping approach 96%
- Genetic modulation of yield and phenotypic plasticity of yield in winter wheat 94%
- The double round-robin population unravels the genetic architecture of grain size in barley 94%
"Similar papers" are the closest papers from that journal in the model's embedding space. They show what the match is built on, but the ranking comes mostly from a classifier over the whole training set, not from these examples alone.