Back

Language Model Embedding Classifiers Enable Identification of Multiple Sclerosis-Associated BCRs and Repertoires

Peet, G. C.; Owens, G. P.; Bennett, J. L.; Krishnan, A.; Macklin, W. B.

2026-07-13 bioinformatics
10.64898/2026.07.07.735316 bioRxiv
Show abstract

Multiple sclerosis (MS) is a chronic inflammatory demyelinating disease. It affects over 2 million people worldwide but has historically been challenging to diagnose, categorize, and treat. MS has an autoimmune component that involves the production of B-cell receptors (BCRs) and immunoglobulins (Igs) that are associated with disease pathophysiology. We reanalyzed all publicly available RNA sequencing data from MS patients and extracted over 11 million BCR immunoglobulin heavy chain (IGH) sequences obtained from a variety of tissue sources. We developed a new decoder-only BCR DNA embedding model that outperforms state-of-the-art BCR and general-purpose protein language models on sequence embedding tasks. We then trained a language model classifier capable of identifying BCR sequences associated with MS. Using low-dimensional representations of whole repertoire embeddings combined with sequence-disease predictions, we can distinguish MS patient repertoires from healthy, infectious disease, or other autoimmune disease repertoires. Our models also successfully rank known MS-associated myelin-binding IgG sequences relative to controls. These findings provide a methodological foundation for BCR-based MS detection and could facilitate identification and study of disease-associated antibodies from blood. SignificanceAutoimmune diseases are difficult to diagnose and treat. Circulating adaptive immune cells contain accessible information about a patients present and past immune experiences. To better use this information in the context of multiple sclerosis, we extracted a comprehensive dataset of B-cell receptor sequences from archival data, developed a new DNA foundation model to embed these sequences, and applied machine learning methods to classify sequence disease association and patient disease state. These findings will advance our ability to model and predict multiple sclerosis and other diseases.

Matching journals

The top 10 journals account for 50% of the predicted probability mass.

50% of probability mass above

"Similar papers" are the closest papers from that journal in the model's embedding space. They show what the match is built on, but the ranking comes mostly from a classifier over the whole training set, not from these examples alone.