A multi-objective based clustering for identifying clonally-related sequences from high-throughput B cell repertoire data
Abdollahi, N.; De Septenville, A. L.; Ripoche, H.; Davi, F.; Silva Bernardes, J.
Show abstract
The adaptive B cell response is driven by the expansion, somatic hypermutation, and selection of B cell clones. A high number of clones in a B cell population indicates a highly diverse repertoire, while clonal size distribution and sequence diversity within clones can be related to antigens selective pressure. Identifying clones is fundamental to many repertoire studies, including repertoire comparisons, clonal tracking and statistical analysis. Several methods have been developed to group sequences from high-throughput B cell repertoire data. Current methods use clustering algorithms to group clonally-related sequences based on their similarities or distances. Such approaches create groups by optimizing a single objective that typically minimizes intra-clonal distances. However, optimizing several objective functions can be advantageous and boost the algorithm convergence rate. Here we propose a new method based on multi-objective clustering. Our approach requires V(D)J annotations to obtain the initial clones and iteratively applies two objective functions that optimize cohesion and separation within clones simultaneously. We show that under simulations with varied mutation rates, our method greatly improves clonal grouping as compared to other tools. When applied to experimental repertoires generated from high-throughput sequencing, its clustering results are comparable to the most performing tools. The method based on multi-objective clustering can accurately identify clone members, has fewer parameter settings and presents the lowest running time among existing tools. All these features constitute an attractive option for repertoire analysis, particularly in the clinical context to unravel the mechanisms involved in the development and evolution of B cell malignancies.
Matching journals
The top 4 journals account for 50% of the predicted probability mass.
Similar papers in this journal
- scConsensus: combining supervised and unsupervised clustering for cell type identification in single-cell RNA sequencing data 94%
- Frugal alignment-free identification of FLT3-internal tandem duplications with FiLT3r 93%
- DAESC+: High-performance, integrated software for single-cell allele-specific expression data 93%
Similar papers in this journal
- Somatic hypermutation analysis for improved identification of B cell clonal families from next-generation sequencing data 97%
- nf-core/airrflow: an adaptive immune receptor repertoire analysis workflow employing the Immcantation framework 95%
- Nucleotide context models outperform protein language models for predicting antibody affinity maturation 94%
Similar papers in this journal
- ViCloD, an interactive web tool for visualizing B cell repertoires and analyzing intra-clonal diversities: application to human B-cell tumors 97%
- Benchmarking computational methods for B-cell receptor reconstruction from single-cell RNA-seq data 96%
- Clone decomposition based on mutation signatures provides novel insights into mutational processes 94%
Similar papers in this journal
- An unbiased comparison of immunoglobulin sequence aligners 96%
- CosTaL: An Accurate and Scalable Graph-Based Clustering Algorithm for High-Dimensional Single-Cell Data Analysis 94%
- Hierarchical cell-type identifier accurately distinguishes immune-cell subtypes enabling precise profiling of tissue microenvironment with single-cell RNA-sequencing 93%
"Similar papers" are the closest papers from that journal in the model's embedding space. They show what the match is built on, but the ranking comes mostly from a classifier over the whole training set, not from these examples alone.