Expanding the Landscape of Disordered Flexible Linkers: A Structural and Computational Framework for DLD dataset assembly
meng, d.; Glavina, J.; Pollastri, G.; Chemes, L. B.
Show abstract
Disordered flexible linkers (DFLs) are functional elements found within intrinsically disordered regions that carry out key functions by connecting domains and/or short linear motifs. Understanding the features of DFLs is limited by the lack of comprehensive datasets and accurate predictive models. In this study, we propose a classification for DFLs that includes linkers joining two domains (DLD), a domain and a motif (MLD/DLM) or two short linear motifs (MLM). We developed a workflow that allows the systematic identification of DLD-type linkers from protein structures and created a comprehensive dataset known as the DLD dataset. The DLD dataset includes 1640 independent domain linkers (IDLs) which significantly expands all currently available linker datasets and additionally annotates related regions such as dependent-domain linkers, intra-domain loops, and termini. Our data collection process integrates missing residue completion and smoothing of short secondary structure stretches enabling to capture a higher number of longer IDLs. We assessed the features of IDLs using t-SNE analysis and protein language model embedding with a CNN-based classifier. IDLs can be distinguished from other disordered and folded protein regions, and their features highly overlap with DisProt Linkers, considered the gold standard for linker annotation. The DLD dataset offers a valuable resource for researchers seeking to investigate the features of disordered flexible linkers and to improve the accuracy and generalizability of DFL predictive models. The methodology is scalable and can be easily applied to larger datasets of protein structures. The DLD dataset is available at https://pcrgwd.ucd.ie/linker/.
Matching journals
The top 4 journals account for 50% of the predicted probability mass.
Similar papers in this journal
"Similar papers" are the closest papers from that journal in the model's embedding space. They show what the match is built on, but the ranking comes mostly from a classifier over the whole training set, not from these examples alone.