Back

Multiscale Temporal Processing Supports Sound Recognition under Causal Constraints

Esposito, M.; Weidler, T.; Ferreyra, C.; Giordano, B. L.; Formisano, E.

2026-07-30 neuroscience
10.64898/2026.07.27.740874 bioRxiv
Show abstract

The auditory system operates under a fundamental computational constraint: at any moment, it has access only to past and present acoustic information. At the same time, it processes sounds across multiple temporal scales, although the computational advantage of this organization remains unclear. Despite both being defining characteristics of biological auditory processing, causal processing and multiscale temporal processing are rarely considered together in computational models. Here, we introduce the Multiscale Convolutional Recurrent Neural Network (MSCRNN) architecture, a brain-inspired model designed to test whether processing sounds across multiple temporal contexts improves recognition under causal constraints. Multiscale processing substantially improved recognition under causal constraints, enabling performance comparable to non-causal architectures after an initial evidence-accumulation period while providing only limited benefit when future acoustic information was available. Analyses of the networks internal representations revealed a progressive transformation from stream-specific acoustic representations to increasingly integrated category-level representations across recurrent processing stages. Together, these findings provide a computational rationale for why biological auditory systems may benefit from multiscale temporal processing under causal constraints and establish the MSCRNN as an interpretable framework for generating and testing hypotheses about the temporal dynamics of auditory processing using electrophysiological and neuroimaging data.

Matching journals

The top 6 journals account for 50% of the predicted probability mass.

50% of probability mass above

"Similar papers" are the closest papers from that journal in the model's embedding space. They show what the match is built on, but the ranking comes mostly from a classifier over the whole training set, not from these examples alone.