Back

Lisen&Curate: A platform to facilitate knowledge tools for curation of regulation of transcription initiation in bacteria

Mendez-Cruz, C.-F.; Diaz-Rodriguez, M.; Guadarrama-Garcia, F.; Lithgow-Serrano, O. W.; Gama-Castro, S.; Solano-Lira, H.; Rinaldi, F.; Collado-Vides, J.

2020-09-23 bioinformatics
10.1101/2020.04.28.065243 bioRxiv
Show abstract

The amount of published papers in biomedical research makes it rather impossible for a researcher to keep up to date. This is where machine processing of scientific publications could contribute to facilitate the access to knowledge. How to make use of text mining capabilities and still preserve the high quality of manual curation, is the challenge we focused on. Here we present the Lisen&Curate system designed to enable current and future NLP capabilities within a curation environment interface used in curation of literature on the regulation of transcription initiation in bacteria. The current version extracts regulatory interactions with the corresponding sentences for curators to confirm or reject accelerating their curation. It also uses an embedded metrics of sentence similarity offering the curator an alternative mechanism of navigating through semantically similar sentences within a given paper as well as across papers of a pre-defined corpus of publications pertinent to the task. We show results of the use of the system to curate literature in E. coli as well as literature in Salmonella. A major advantage of the system is to save as part of the curation work, the precise link for every curated piece of knowledge with the corresponding specific sentence(s) in the curated publication supporting it. We discuss future directions of this type of curation infrastructure.

Matching journals

The top 3 journals account for 50% of the predicted probability mass.

1
Database
61 papers in training set
Top 0.1%
26.8%
2
Nucleic Acids Research
1281 papers in training set
Top 1%
13.0%
3
Bioinformatics
1204 papers in training set
Top 2%
12.0%
50% of probability mass above
4
BMC Bioinformatics
457 papers in training set
Top 0.9%
8.0%
5
GigaScience
212 papers in training set
Top 1%
3.5%
6
PLOS ONE
5266 papers in training set
Top 37%
3.3%
7
NAR Genomics and Bioinformatics
242 papers in training set
Top 2%
2.8%
8
PLOS Computational Biology
1863 papers in training set
Top 12%
2.4%
9
Bioinformatics Advances
203 papers in training set
Top 3%
1.9%
10
F1000Research
88 papers in training set
Top 1%
1.7%
11
Computational and Structural Biotechnology Journal
242 papers in training set
Top 4%
1.5%
12
Scientific Data
209 papers in training set
Top 2%
1.1%
13
BMC Genomics
406 papers in training set
Top 6%
1.1%
14
PeerJ
308 papers in training set
Top 8%
1.1%
15
Frontiers in Bioinformatics
49 papers in training set
Top 1%
1.0%
16
Briefings in Bioinformatics
354 papers in training set
Top 6%
1.0%
17
eLife
5828 papers in training set
Top 60%
1.0%
18
Journal of Molecular Biology
232 papers in training set
Top 3%
0.9%
19
G3 Genes|Genomes|Genetics
351 papers in training set
Top 4%
0.9%
20
Gigabyte
62 papers in training set
Top 1%
0.9%
21
BMC Biology
265 papers in training set
Top 5%
0.9%
22
Scientific Reports
3612 papers in training set
Top 73%
0.9%
23
BMC Research Notes
33 papers in training set
Top 2%
0.6%
24
Frontiers in Genetics
230 papers in training set
Top 6%
0.6%