Combined Linkage Disequilibrium and Linkage Analysis (cLDLA): implementation of a powerful approach to identify the genetic basis of complex traits in a bioinformatics workflow.
Upadhyay, M.; Panchal, A.; Medugorac, I.
Show abstract
Identifying the relationship between the polymorphism segregating in a population and phenotypic differences of a trait observed between the individuals of a population is of major biological interest and represents the basis of forward genetics. Much of the traits of interest are influenced by several polymorphic genes and environmental conditions. Often the loci associated with such measurable traits are referred to as Quantitative trait loci. These loci are identified using several statistical approaches. One of them is combined linkage disequilibrium and linkage analysis (cLDLA). This approach, first proposed by Meuwissen and colleagues in 2002, is shown to be robust against population stratification/family structure and requires a relatively lower sample size compared to a genome-wide association study design. Previously, we have successfully used this approach in mapping several important traits in livestock such as identifying the genetic basis of polled condition in cattle and tail length in sheep. A cLDLA requires several complex computation processing and intermediary file conversion steps; for some of these steps no open-source tools are available. Therefore, running this analysis, manually, can prove challenging, tedious, or error-prone. We present, cldla, a bioinformatics workflow implemented in nextflow which takes the vcf file and phenotype file as inputs and implements all the downstream processing required for cLDLA. Additionally, it also has a separate workflow to estimate SNP-based heritability and features for interactive visualization of the results. The workflow is freely available at: https://github.com/Popgen48/cldla.
Matching journals
The top 5 journals account for 50% of the predicted probability mass.
Similar papers in this journal
- BWGS: a R package for genomic selection and its application to a wheat breeding programme. 96%
- In-situ genomic prediction using low-coverage Nanopore sequencing 95%
- NGSpop: A desktop software that supports population studies by identifying sequence variations from next-generation sequencing data 94%
Similar papers in this journal
- Dimensionality of genomic information and its impact on GWA and variant selection: a simulation study 95%
- Bayesian genomic models boost prediction accuracy for resistance against Streptococcus agalactiae in Nile tilapia (Oreochromus nilioticus) 95%
- Optimisation of the core subset for the APY approximation of genomic relationships 94%
Similar papers in this journal
- A Nextflow pipeline for molecular quantitative trait loci mapping in small sample size datasets with an application in Atlantic salmon 94%
- RecView: an interactive R application for viewing and locating recombination positions using pedigree data 94%
- Benchmarking phasing software with a whole-genome sequenced cattle pedigree 94%
"Similar papers" are the closest papers from that journal in the model's embedding space. They show what the match is built on, but the ranking comes mostly from a classifier over the whole training set, not from these examples alone.