intern: Integrated Toolkit for Extensible and Reproducible Neuroscience
Matelsky, J. K.; Rodriguez, L.; Xenes, D.; Gion, T.; Hider, R.; Wester, B.; Gray-Roncal, W.
Show abstract
As neuroscience datasets continue to grow in size, the complexity of data analyses can require a detailed understanding and implementation of systems computer science for storage, access, processing, and sharing. Currently, several general data standards (e.g., Zarr, HDF5, precompute, tensorstore) and purpose-built ecosystems (e.g., BossDB, CloudVolume, DVID, and Knossos) exist. Each of these systems has advantages and limitations and is most appropriate for different use cases. Using datasets that dont fit into RAM in this heterogeneous environment is challenging, and significant barriers exist to leverage underlying research investments. In this manuscript, we outline our perspective for how to approach this challenge through the use of community provided, standardized interfaces that unify various computational backends and abstract computer science challenges from the scientist. We introduce desirable design patterns and our reference implementation called intern.
Matching journals
The top 6 journals account for 50% of the predicted probability mass.
Similar papers in this journal
- Datavzrd: Rapid programming- and maintenance-free interactive visualization and communication of tabular data 96%
- Advancing clinical cohort selection with genomics analysis on a distributed platform 96%
- Sardine: a modular framework for developing data acquisition and near real-time analysis applications 94%
"Similar papers" are the closest papers from that journal in the model's embedding space. They show what the match is built on, but the ranking comes mostly from a classifier over the whole training set, not from these examples alone.