Back

The ground truth of the Data-Iceberg: Correct Meta-data

Caliskan, A.; Dangwal, S.; Dandekar, T.

2021-12-21 bioinformatics
10.1101/2021.12.17.473021 bioRxiv
Show abstract

Short summaryBiological molecular data such as sequence information increase so rapidly that detailed metadata, describing the process and conditions of data collection as well as proper labelling and typing of the data become ever more important to avoid mistakes and erroneous labeling. Starting from a striking example of wrong labelling of patient data recently published in Nature, we advocate measures to improve software metadata and controls in a timely manner to not rapidly loose quality in the ever-growing data flood.

Matching journals

The top 11 journals account for 50% of the predicted probability mass.

50% of probability mass above

"Similar papers" are the closest papers from that journal in the model's embedding space. They show what the match is built on, but the ranking comes mostly from a classifier over the whole training set, not from these examples alone.