Back

Note: Updating the metadata of four misidentified samples in the DrosRTEC dataset

Nunez, J. C. B.; Paris, M.; Machado, H.; Bogaerts, M.; Gonzalez, J.; Flatt, T.; Coronado, M.; Kapun, M.; Schmidt, P.; Petrov, D.; Bergland, A.

2021-01-27 genomics
10.1101/2021.01.26.428249 bioRxiv
Show abstract

This note details the consortiums rationale behind its decision to modify the metadata for putatively misidentified European samples in the DrosRTEC dataset. In brief, we use PCA on published datasets from North America and Europe to generate phylogeographic clusters reflective of worldwide D. melanogaster demography. We used this PCA to train a DAPC model in order the predict the group membership. Our results indicate that 4 out of 73 samples were misclassified and the metadata was updated accordingly. These samples are a spring-fall pair from Spain and Austria.

Matching journals

The top 5 journals account for 50% of the predicted probability mass.

50% of probability mass above

"Similar papers" are the closest papers from that journal in the model's embedding space. They show what the match is built on, but the ranking comes mostly from a classifier over the whole training set, not from these examples alone.