Back

Scalable graph analysis tools for the connectomics community

Matelsky, J. K.; Johnson, E. C.; Wester, B.; Gray-Roncal, W. R.

2022-06-02 neuroscience
10.1101/2022.06.01.494307 bioRxiv
Show abstract

Neuroscientists now have the opportunity to analyze synaptic resolution connectomes that are larger than the memory on single consumer workstations. As dataset size and tissue diversity have grown, there is increasing interest in conducting comparative connectomics research, including rapidly querying and searching for recurring patterns of connectivity across brain regions and species. There is also a demand for algorithm reuse -- applying methods developed for one dataset to another volume. A key technological hurdle is enabling researchers to efficiently and effectively query these diverse datasets, especially as the raw image volumes grow beyond terabyte sizes. Existing community tools can perform such queries and analysis on smaller scale datasets, which can fit locally in memory, but the path to scaling remains unclear. Existing solutions such as neuPrint or FlyBrainLab enable these queries for specific datasets, but there remains a need to generalize algorithms and standards across datasets. To overcome this challenge, we present a software framework for comparative connectomics and graph discovery to make connectomes easy to analyze, even when larger-than-RAM, and even when stored in disparate datastores. This software suite includes visualization tools, a web portal, a connectivity and annotation query engine, and the ability to interface with a variety of data sources and community tools from the neuroscience community. These tools include MossDB (an immutable datastore for metadata and rich annotations); Grand (for prototyping larger-than-RAM graphs); GrandIso-Cloud (for querying existing graphs that exceed the capabilities of a single work-station); and Motif Studio (for enabling the public to query across connectomes). These tools interface with existing frameworks such as neuPrint, graph databases such as Neo4j, and standard data analysis tools such as Pandas or NetworkX. Together, these tools enable tool and algorithm reuse, standardization, and neuroscience discovery.

Matching journals

The top 6 journals account for 50% of the predicted probability mass.

1
PLOS Computational Biology
1863 papers in training set
Top 1%
18.5%
2
Journal of Open Source Software
25 papers in training set
Top 0.1%
15.1%
3
Nature Methods
385 papers in training set
Top 1%
6.7%
4
Nature Computational Science
55 papers in training set
Top 0.1%
5.5%
5
Bioinformatics
1204 papers in training set
Top 5%
4.0%
6
Patterns
78 papers in training set
Top 0.5%
3.4%
50% of probability mass above
7
BMC Bioinformatics
457 papers in training set
Top 3%
3.2%
8
PLOS ONE
5266 papers in training set
Top 37%
3.2%
9
Nucleic Acids Research
1281 papers in training set
Top 6%
2.8%
10
eLife
5828 papers in training set
Top 37%
2.8%
11
GigaScience
212 papers in training set
Top 1%
2.8%
12
SoftwareX
15 papers in training set
Top 0.1%
2.6%
13
eneuro
439 papers in training set
Top 3%
2.4%
14
Genome Biology
637 papers in training set
Top 6%
1.3%
15
NAR Genomics and Bioinformatics
242 papers in training set
Top 3%
1.3%
16
Cell Reports Methods
165 papers in training set
Top 2%
1.3%
17
Proceedings of the National Academy of Sciences
2444 papers in training set
Top 35%
1.1%
18
Frontiers in Neuroinformatics
41 papers in training set
Top 0.5%
1.1%
19
F1000Research
88 papers in training set
Top 2%
1.1%
20
Bioinformatics Advances
203 papers in training set
Top 4%
1.1%
21
G3 Genes|Genomes|Genetics
351 papers in training set
Top 3%
1.1%
22
Nature Communications
5641 papers in training set
Top 54%
1.0%
23
BMC Genomics
406 papers in training set
Top 8%
0.8%
24
STAR Protocols
18 papers in training set
Top 0.3%
0.8%
25
Nature Biotechnology
172 papers in training set
Top 4%
0.8%
26
Journal of Proteome Research
234 papers in training set
Top 2%
0.8%
27
Neuron
337 papers in training set
Top 5%
0.6%
28
Cell Genomics
172 papers in training set
Top 4%
0.6%
29
Nature Protocols
33 papers in training set
Top 0.6%
0.6%