Back

Evaluation of Artificial Intelligence (AI)-based in silico tools for variant classification in clinically actionable NSCLC variants

Kim, E.; Duarte, S. E.; Yu, E.; Hong, I.; Lee, G.; Chae, Y. K.

2024-04-14 oncology
10.1101/2024.04.12.24305738 medRxiv
Show abstract

Introduction/BackgroundGenetic variants beyond FDA-approved drug targets are often identified in non-small cell lung cancer (NSCLC) patients. Although the performances of in silico tools in predicting variant pathogenicity have been analyzed in previous studies, they have not been analyzed for actionable targets of FDA-approved therapies for NSCLC. The aim of this study is to compare the performance of commonly used in silico tools in classifying the pathogenicity of actionable variants in NSCLC. Materials and MethodsWe evaluated the performance of the following in silico tools: Polyphen-2 (HumDiv, HumVar), Align-GVGD, MutationTaster2021, CADD, CONDEL, and REVEL. A curated set of targetable NSCLC missense variants (n=236) was used. The overall accuracy, sensitivity, specificity, positive predictive value (PPV), negative predictive value (NPV), and Matthews correlation coefficient (MCC) of each in silico tool was determined. ResultsThe most recently released MutationTaster2021 demonstrated the highest performance in terms of accuracy, specificity, PPV, and MCC, but was outperformed by CADD for both sensitivity and NPV. Although some tools demonstrated high sensitivities, all tools except MutationTaster2021 displayed markedly low overall specificities, as low as 23%. ConclusionThe collective results indicate that the evaluated in silico tools can provide guidance in predicting the pathogenicity of NSCLC missense variants, but are not fully reliable. The tools analyzed in this study could be acceptable to rule out pathogenicity in variants given their higher sensitivities, but are limited when it comes to identifying pathogenicity in variants due to low specificities. HighlightsO_LIMutationTaster2021 demonstrated the highest overall performance C_LIO_LICADD demonstrated the highest sensitivity (99.19%) and NPV (96.55%). C_LIO_LIAll tools but MutationTaster2021 demonstrated low specificities, as low as 23% C_LIO_LIAn individual predictor outperformed 3 meta predictors C_LIO_LIPerformance variability suggests caution when using in silico tools in the clinic C_LI

Matching journals

The top 8 journals account for 50% of the predicted probability mass.

50% of probability mass above

"Similar papers" are the closest papers from that journal in the model's embedding space. They show what the match is built on, but the ranking comes mostly from a classifier over the whole training set, not from these examples alone.