Back

intern: Integrated Toolkit for Extensible and Reproducible Neuroscience

Matelsky, J. K.; Rodriguez, L.; Xenes, D.; Gion, T.; Hider, R.; Wester, B.; Gray-Roncal, W.

2020-05-16 neuroscience
10.1101/2020.05.15.098707 bioRxiv
Show abstract

As neuroscience datasets continue to grow in size, the complexity of data analyses can require a detailed understanding and implementation of systems computer science for storage, access, processing, and sharing. Currently, several general data standards (e.g., Zarr, HDF5, precompute, tensorstore) and purpose-built ecosystems (e.g., BossDB, CloudVolume, DVID, and Knossos) exist. Each of these systems has advantages and limitations and is most appropriate for different use cases. Using datasets that dont fit into RAM in this heterogeneous environment is challenging, and significant barriers exist to leverage underlying research investments. In this manuscript, we outline our perspective for how to approach this challenge through the use of community provided, standardized interfaces that unify various computational backends and abstract computer science challenges from the scientist. We introduce desirable design patterns and our reference implementation called intern.

Matching journals

The top 6 journals account for 50% of the predicted probability mass.

50% of probability mass above

"Similar papers" are the closest papers from that journal in the model's embedding space. They show what the match is built on, but the ranking comes mostly from a classifier over the whole training set, not from these examples alone.