Back

Leveraging big data for classification of children who stutter from fluent peers

Rutherford, S.; Angstadt, M.; Sripada, C.; Chang, S.-E.

2020-10-29 neuroscience
10.1101/2020.10.28.359711 bioRxiv
Show abstract

IntroductionLarge datasets, consisting of hundreds or thousands of subjects, are becoming the new data standard within the neuroimaging community. While big data creates numerous benefits, such as detecting smaller effects, many of these big datasets have focused on non-clinical populations. The heterogeneity of clinical populations makes creating datasets of equal size and quality more challenging. There is a need for methods to connect these robust large datasets with the carefully curated clinical datasets collected over the past decades. MethodsIn this study, resting-state fMRI data from the Adolescent Brain Cognitive Development study (N=1509) and the Human Connectome Project (N=910) is used to discover generalizable brain features for use in an out-of-sample (N=121) multivariate predictive model to classify young (3-10yrs) children who stutter from fluent peers. ResultsAccuracy up to 72% classification is achieved using 10-fold cross validation. This study suggests that big data has the potential to yield generalizable biomarkers that are clinically meaningful. Specifically, this is the first study to demonstrate that big data-derived brain features can differentiate children who stutter from their fluent peers and provide novel information on brain networks relevant to stuttering pathophysiology. DiscussionThe results provide a significant expansion to previous understanding of the neural bases of stuttering. In addition to auditory, somatomotor, and subcortical networks, the big data-based models highlight the importance of considering large scale brain networks supporting error sensitivity, attention, cognitive control, and emotion regulation/self-inspection in the neural bases of stuttering.

Matching journals

The top 7 journals account for 50% of the predicted probability mass.

50% of probability mass above

"Similar papers" are the closest papers from that journal in the model's embedding space. They show what the match is built on, but the ranking comes mostly from a classifier over the whole training set, not from these examples alone.