Back

Negative Binomial Mixture Model for Identification of Noise in Antigen-Specificity Predictions by LIBRA-seq

Wasdin, P. T.; Abu-Shmais, A. A.; Irvin, M. W.; Vukovich, M. J.; Georgiev, I. S.

2023-10-17 bioinformatics
10.1101/2023.10.13.562258 bioRxiv
Show abstract

Structured AbstractO_ST_ABSMotivationC_ST_ABSLIBRA-seq (linking B cell receptor to antigen specificity by sequencing) provides a powerful tool for interrogating the antigen-specific B cell compartment and identifying antibodies against antigen targets of interest. Identification of noise in LIBRA-seq antigen count data is critical for improving antigen binding predictions for downstream applications including antibody discovery and machine learning technologies. ResultsIn this study, we present a method for denoising LIBRA-seq data by clustering antigen counts into signal and noise components with a negative binomial mixture model. This approach leverages the VRC01 negative control cells included in a recent LIBRA-seq study(Abu-Shmais et al.) to provide a data-driven means for identification of technical noise. We apply this method to a dataset of nine donors representing separate LIBRA-seq experiments and show that our approach provides improved predictions for in vitro antibody-antigen binding when compared to the standard scoring method used in LIBRA-seq, despite variance in data size and noise structure across samples. This development will improve the ability of LIBRA-seq to identify antigen-specific B cells and contribute to providing more reliable datasets for future machine learning based approaches to predicting antibody-antigen binding as the corpus of LIBRA-seq data continues to grow. Availability and ImplementationJupyter notebooks detailing model fitting and figure generation in Python are available at https://github.com/perrywasdin/mixture_model_denoising. ContactEmail: Ivelin.Georgiev@Vanderbilt.edu Supplementary InformationSupplementary figures are provided in the attached PDF.

Matching journals

The top 7 journals account for 50% of the predicted probability mass.

50% of probability mass above

"Similar papers" are the closest papers from that journal in the model's embedding space. They show what the match is built on, but the ranking comes mostly from a classifier over the whole training set, not from these examples alone.