Back

LOCC: a novel visualization and scoring of cutoffs for continuous variables

Luo, G.; Letterio, J.

2023-04-12 bioinformatics
10.1101/2023.04.11.536461 bioRxiv
Show abstract

ObjectiveThere is a need for new methods to select and analyze cutoffs employed to define genes that are most prognostic significant and impactful. We designed LOCC (Luos Optimization Categorization Curve), a novel tool to visualize and score continuous variables for a dichotomous outcome. MethodsTo demonstrate LOCC with real world data, we analyzed TCGA hepatocellular carcinoma gene expression and patient data using LOCC. We compared LOCC visualization to receiver operating characteristic (ROC) curve for prognostic modeling to showcase its utility in understanding predictors in various TCGA datasets. ResultsAnalysis of E2F1 expression in hepatocellular carcinoma using LOCC demonstrated appropriate cutoff selection and validation. In addition, we compared LOCC visualization and scoring to ROC curves and c-statistics, demonstrating that LOCC better described predictors. Analysis of a previously published gene signature showed large differences in LOCC scoring, and removing the lowest scoring genes did not affect prognostic modeling of the gene signature demonstrating LOCC scoring could distinguish which predictors were most critical. ConclusionOverall, LOCC is a novel visualization tool for understanding and selecting cutoffs, particularly for gene expression analysis in cancer. The LOCC score can be used to rank genes for prognostic potential and is more suitable than ROC curves for prognostic modeling. Graphical Abstract O_FIG O_LINKSMALLFIG WIDTH=200 HEIGHT=95 SRC="FIGDIR/small/536461v2_ufig1.gif" ALT="Figure 1"> View larger version (17K): org.highwire.dtl.DTLVardef@1712bc2org.highwire.dtl.DTLVardef@efd579org.highwire.dtl.DTLVardef@1a8236borg.highwire.dtl.DTLVardef@1ad78ea_HPS_FORMAT_FIGEXP M_FIG C_FIG

Matching journals

The top 11 journals account for 50% of the predicted probability mass.

50% of probability mass above

"Similar papers" are the closest papers from that journal in the model's embedding space. They show what the match is built on, but the ranking comes mostly from a classifier over the whole training set, not from these examples alone.