Back

GENVISAGE: Rapid Identification of Discriminative and Explainable Feature Pairs for Genomic Analysis

Huang, S.; Blatti, C.; Sinha, S.; Parameswaran, A.

2020-02-05 bioinformatics
10.1101/2020.02.05.935411 bioRxiv
Show abstract

MotivationA common but critical task in genomic data analysis is finding features that separate and thereby help explain differences between two classes of biological objects, e.g., genes that explain the differences between healthy and diseased patients. As lower-cost, high-throughput experimental methods greatly increase the number of samples that are assayed as objects for analysis, computational methods are needed to quickly provide insights into high-dimensional datasets with tens of thousands of objects and features. ResultsWe develop an interactive exploration tool called GO_SCPLOWENVISAGEC_SCPLOW that rapidly discovers the most discriminative feature pairs that best separate two classes in a dataset, and displays the corresponding visualizations. Since quickly finding top feature pairs is computationally challenging, especially when the numbers of objects and features are large, we propose a suite of optimizations to make GO_SCPLOWENVISAGEC_SCPLOW more responsive and demonstrate that our optimizations lead to a 400X speedup over competitive baselines for multiple biological data sets. With this speedup, GO_SCPLOWENVISAGEC_SCPLOW enables the exploration of more large-scale datasets and alternate hypotheses in an interactive and interpretable fashion. We apply GO_SCPLOWENVISAGEC_SCPLOW to uncover pairs of genes whose transcriptomic responses significantly discriminate treatments of several chemotherapy drugs. AvailabilityFree webserver at http://genvisage.knoweng.org:443/ with source code at https://github.com/KnowEnG/Genvisage

Matching journals

The top 5 journals account for 50% of the predicted probability mass.

50% of probability mass above

"Similar papers" are the closest papers from that journal in the model's embedding space. They show what the match is built on, but the ranking comes mostly from a classifier over the whole training set, not from these examples alone.