OmniExtract: An automatic data extraction tool based on Large Language Model and Prompt Engineering
Wang, Y.; Tang, B.; Wu, S.; Meng, Y.; Kong, D.; Zhao, W.
Show abstract
Extracting structured information from documents or scientific papers is crucial for data sharing and retrieval. Recently, Large Language Model (LLM) has shown its impressive ability in text understanding and several tools based on LLM has been developed. However, its still difficult to find a universal and user-friendly tool for various practical extraction tasks. To address this challenge, we propose OmniExtract, an automatic data extraction tool with user-friendly configuration files which can adapt to various data extraction tasks. OmniExtract uses a prompt optimized engineering to improve prompt and obtain high performance, and it can support a comprehensive data extraction including text and tables. Evaluation results show that OmniExtract obtains a high accuracy over 80% for 3 datasets. Furthermore, two additional data extraction applications using OmniExtract have been provided, achieving an accuracy of 92.21% and an average F1 score of 0.83 respectively. The data reliability performance shows that OmniExtract is a valuable tool for database updating.
Matching journals
The top 7 journals account for 50% of the predicted probability mass.
Similar papers in this journal
Similar papers in this journal
Similar papers in this journal
Similar papers in this journal
- Academic Tracker: Software for Tracking and Reporting Publications Associated with Authors and Grants 95%
- Datavzrd: Rapid programming- and maintenance-free interactive visualization and communication of tabular data 94%
- Extract, Transform, Load Framework for the Conversion of Health Databases to OMOP 94%
Similar papers in this journal
- Optimizing biomedical information retrieval with a keyword frequency-driven Prompt Enhancement Strategy 96%
- Relation extraction between bacteria and biotopes from biomedical texts with attention mechanisms and domain-specific contextual representations 94%
- Transformer-based tool recommendation system in Galaxy 94%
"Similar papers" are the closest papers from that journal in the model's embedding space. They show what the match is built on, but the ranking comes mostly from a classifier over the whole training set, not from these examples alone.