Quantifying Data Distortion in Bar Graphs in Biological Research
Lin, T.-J.; Landry, M. P.
10.1101/2024.09.20.609464 bioRxivShow abstract
Over 88% of biological research articles use bar graphs, of which 29% have undocumented data distortion mistakes that over- or under-state findings. We developed a framework to quantify data distortion and analyzed bar graphs published across 3387 articles in 15 journals, finding consistent data distortions across journals and common biological data types. To reduce bar graph-induced data distortion, we propose recommendations to improve data visualization literacy and guidelines for effective data visualization.
Matching journals
The top 3 journals account for 50% of the predicted probability mass.
Similar papers in this journal
- Density-Preserving Data Visualization Unveils Dynamic Patterns of Single-Cell Transcriptomic Variability 94%
- Multi-omics integration and regulatory inference for unpaired single-cell data with a graph-linked unified embedding framework 94%
- Comparative analysis of cell-cell communication at single-cell resolution 94%
Similar papers in this journal
- Single-cell epigenomic reconstruction of developmental trajectories in human neural organoid systems from pluripotency 94%
- Protosequences in brain organoids model intrinsic brain states 93%
- A deep learning framework for inference of single-trial neural population dynamics from calcium imaging with sub-frame temporal resolution 92%
Similar papers in this journal
Similar papers in this journal
- Glial place cells: complementary encoding of spatial information in hippocampal astrocytes 93%
- A unifying framework disentangles genetic, epigenetic, and stochastic sources of drug-response variability in an in vitro model of tumor heterogeneity 93%
- The X-linked splicing regulator MBNL3 has been co-opted to restrict placental growth in eutherians 92%
"Similar papers" are the closest papers from that journal in the model's embedding space. They show what the match is built on, but the ranking comes mostly from a classifier over the whole training set, not from these examples alone.