Back

Novel Protein Structure Validation using PDBMine and Data Analytics Approaches

Pandala, N.; Brown, K. G.; Valafar, H.

2025-02-08 bioinformatics
10.1101/2025.02.07.637116 bioRxiv
Show abstract

Protein structure prediction is essential for understanding biological functions and advancing drug development. Although experimental techniques like NMR, X-ray crystallography, and cryo-EM provide valuable insights, they are expensive and time-consuming, prompting reliance on computational approaches. AlphaFold2 revolutionized protein model predictions accuracy in 2020. However, limitations remain in the prediction of novel proteins, complex conformations, and mutations. To address these challenges, we leverage PDBMine software and machine learning for advanced data analytics. This approach detects and corrects structural inaccuracies, calculates fitness scores, and enhances model reliability, accelerating drug discovery and therapeutic breakthroughs by bridging gaps in current protein prediction capabilities.

Matching journals

The top 6 journals account for 50% of the predicted probability mass.

50% of probability mass above

"Similar papers" are the closest papers from that journal in the model's embedding space. They show what the match is built on, but the ranking comes mostly from a classifier over the whole training set, not from these examples alone.