Back

variant2literature: full text literature search for genetic variants

Lin, Y.-H.; Lu, Y.-C.; Chen, T.-F.; Hsu, J. S.; Lee, K.-H.; Cheng, Y.-W.; Chen, Y.-C.; Fan, J.-S.; Tu, C.-T.; Hsu, C.-M.; Chou, C.-C.; Chen, P.-L.; Tu, Y.-C. E.; Chen, C.-Y.

2019-06-04 genetics
10.1101/583450 bioRxiv
Show abstract

MotivationWhole genome sequencing (WGS) by next-generation sequencing produces millions of variants for an individual. The retrieval of biomedical literature for such a large number of genetic variants remains challenging, because in many cases the variants are only present in tables as images, or in the supplementary documents of which the file formats are diverse.\n\nResultsThe proposed tool named variant2literature from the TaiGenomics (Toolkits for AI genomics) resolves the problem by incorporating text recognition with image processing. In addition to the adoption of advanced image-based text retrieval, the recall rate of finding the literature containing the variants of interest is further improved by employing the skill of variant normalization. Different variant presentations are transformed into chromosome coordinates (standard VCF format) such that false negatives can be largely avoided. variant2literature is available in two ways. First, a web-based interface is provided to search all the literature in PMC Open Access Subset. Second, the command-line executable can be downloaded such that the users are free to search all the files in a specified directory locally.\n\nAvailabilityhttp://variant2literature.taigenomics.com/\n\nContactchienyuchen@ntu.edu.tw

Matching journals

The top 9 journals account for 50% of the predicted probability mass.

50% of probability mass above

"Similar papers" are the closest papers from that journal in the model's embedding space. They show what the match is built on, but the ranking comes mostly from a classifier over the whole training set, not from these examples alone.