Enriching Biomedical Knowledge for Low-resource Language Through Translation
Phan, L.; Dang, T.; Tran, H.; Phan, V.; Chau, L. D.; Trinh, T. H.
Show abstract
Biomedical data and benchmarks are highly valuable yet very limited in low-resource languages other than English such as Vietnamese. In this paper, we make use of a state-of-theart translation model in English-Vietnamese to translate and produce both pretrained as well as supervised data in the biomedical domains. Thanks to such large-scale translation, we introduce ViPubmedT5, a pretrained Encoder-Decoder Transformer model trained on 20 million translated abstracts from the high-quality public PubMed corpus. ViPubMedT5 demonstrates state-of-the-art results on two different biomedical benchmarks in summarization and acronym disambiguation. Further, we release ViMedNLI a new NLP task in Vietnamese translated from MedNLI using the recently public En-vi translation model and carefully refined by human experts, with evaluations of existing methods against ViPubmedT5.
Matching journals
The top 4 journals account for 50% of the predicted probability mass.
Similar papers in this journal
Similar papers in this journal
Similar papers in this journal
- A Sequence Labeling Framework for Extracting Drug-Protein Relations from Biomedical Literature 93%
- LSD600: the first corpus of biomedical abstracts annotated with lifestyle–disease relations 92%
- RegulaTome: a corpus of typed, directed, and signed relations between biomedical entities in the scientific literature 91%
"Similar papers" are the closest papers from that journal in the model's embedding space. They show what the match is built on, but the ranking comes mostly from a classifier over the whole training set, not from these examples alone.