Back

Trends in Hearing

SAGE Publications

All preprints, ranked by how well they match Trends in Hearing's content profile, based on 15 papers previously published here. The average preprint has a 0.01% match score for this journal, so anything above that is already an above-average fit. Older preprints may already have been published elsewhere.

1
Spatial auditory change detection in listeners with hearing loss

Poole, K. C.; With, S.; Martin, V.; Chait, M.; Picinali, L.; Shiell, M. M.

2025-12-12 neuroscience 10.64898/2025.12.10.693419 medRxiv
Top 0.1%
31.2%
Show abstract

Everyday listening relies on the auditory systems ability to automatically monitor the background soundscape and detect new or changing sources. Although change detection is a fundamental aspect of situational awareness, little is known about how hearing impairment affects this ability. This study examined how sensorineural hearing loss influences spatial auditory change detection. Older hearing-impaired listeners (N = 30) completed a spatial change detection task requiring them to identify the appearance of a new sound source within a complex spatialised acoustic scene. Hearing loss was characterised by three factors that were measured with standard clinical tests: audiometric hearing thresholds, sensitivity to small level changes, and sensitivity to spectrotemporal modulation. Simple and mixed-effects linear models were used to test how these factors predicted reaction time, hit rate, and false alarm rate. Listeners with poorer spectrotemporal sensitivity, higher audiometric hearing thresholds, and older age showed slower and less accurate detection, whereas sensitivity to small changes in level did not predict outcomes. Detection also varied with spatial location, where appearing sources from behind were detected more slowly and less accurately than those from the front or sides. Numerical analysis using head-related transfer functions confirmed that these rear-field effects were unlikely to be explained by overall or frequency-specific acoustic level differences. These findings reveal that hearing loss, age, and spatial factors jointly shape listeners ability to monitor dynamic auditory scenes. Additionally, testing spectrotemporal sensitivity offers a promising clinical measure of non-speech auditory processing with relevance for hearing-aid fitting and situational awareness.

2
Does the method matter? Evaluating the effectiveness, efficiency and ease of hearing-aid gain self-adjustment

Benecke, J.; Whitmer, W. M.

2026-06-12 otolaryngology 10.64898/2026.06.11.26355463 medRxiv
Top 0.1%
26.9%
Show abstract

In conventional hearing-aid personalisation, clinicians cannot hear what their patients hear, and patients cannot often reliably detect or describe what they hear. Self-adjustment avoids this issue but requires user controls that adjust hearing-aid signal processing parameters to be effective, efficient and easy. In this study, we explored (a) the roles of interface complexity and stimulus type in the self-adjustment of hearing-aid gain, and (b) how well individuals can adjust one sound to match another to assess the same interfaces and stimuli. Adult hearing-aid users with mild to moderate symmetrical sensorineural hearing loss repeatedly adjusted the gain (a) to their preference from individual prescription (n = 41) and (b) to match their previous preferences from a random starting point (n = 32) using three interfaces representing different bass/mid/treble configurations and three stimuli (music, speech and speech-in-noise). The large interindividual variability in self-adjusted gains clustered into three patterns of deviation from initial prescription: increased relative bass, overall gain reduction, and close to initial prescription. There were no substantial effects of interface nor stimulus on self-adjustment reliability (median {sigma} = 2.8 dB), whereas absolute sound-matching error increased with increasing interface complexity and centre frequency. Neither individual matching accuracy nor questionnaire responses predicted either self-adjusted gains or reliability. Overall, these results show that many - but not all - hearing-aid users can adjust gains with reasonable reliability, and while it can be difficult to predict the behaviour from the individual, the individual applies a similar self-adjustment behaviour across different interfaces and stimuli.

3
Individual Differences Reveal the Utility of Temporal Fine-Structure Processing for Speech Perception in Noise

Borjigin, A.; Bharadwaj, H.

2023-09-22 neuroscience 10.1101/2023.09.20.558670 medRxiv
Top 0.1%
26.2%
Show abstract

The auditory system is unique among sensory systems in its ability to phase lock to and precisely follow very fast cycle-by-cycle fluctuations in the phase of sound-driven cochlear vibrations. Yet, the perceptual role of this temporal fine structure (TFS) code is debated. This fundamental gap is attributable to our inability to experimentally manipulate TFS cues without altering other perceptually relevant cues. Here, we circumnavigated this limitation by leveraging individual differences across 200 participants to systematically compare variations in TFS sensitivity to performance in a range of speech perception tasks. TFS sensitivity was assessed through detection of interaural time/phase differences, while speech perception was evaluated by word identification under noise interference. Results suggest that greater TFS sensitivity is not associated with greater masking release from fundamental-frequency or spatial cues, but appears to contribute to resilience against the effects of reverberation. We also found that greater TFS sensitivity is associated with faster response times, indicating reduced listening effort. These findings highlight the perceptual significance of TFS coding for everyday hearing. Significance StatementNeural phase-locking to fast temporal fluctuations in sounds-temporal fine structure (TFS) in particular- is a unique mechanism by which acoustic information is encoded by the auditory system. However, despite decades of intensive research, the perceptual relevance of this metabolically expensive mechanism, especially in challenging listening settings, is debated. Here, we leveraged an individual-difference approach to circumnavigate the limitations plaguing conventional approaches and found that robust TFS sensitivity is associated with greater resilience against the effects of reverberation and is associated with reduced listening effort for speech understanding in noise.

4
A standardised test to evaluate audio-visual speech intelligibility in French

Le Rhun, L.; Llorach, G.; Delmas, T.; Suied, C.; Arnal, L.; Lazard, D.

2023-01-18 otolaryngology 10.1101/2023.01.18.23284110 medRxiv
Top 0.1%
22.0%
Show abstract

ObjectiveLipreading, which plays a major role in the communication of the hearing impaired, lacked a French standardised tool. Our aim was to create and validate an audio-visual (AV) version of the French Matrix Sentence Test (FrMST). DesignVideo recordings were created by dubbing the existing audio files. SampleThirty-five young, normal-hearing participants were tested in auditory and visual modalities alone (Ao, Vo) and in AV conditions, in quiet, noise, and open and closed-set response formats. ResultsLipreading ability (Vo) varied from 1% to 77%-word comprehension. The absolute AV benefit was 9.25[L]dB SPL in quiet and 4.6[L]dB SNR in noise. The response format did not influence the results in the AV noise condition, except during the training phase. Lipreading ability and AV benefit were significantly correlated. ConclusionsThe French video material achieved similar AV benefits as those described in the literature for AV MST in other languages. For clinical purposes, we suggest targeting SRT80 to avoid ceiling effects, and performing two training lists in the AV condition in noise, followed by one AV list in noise, one Ao list in noise and one Vo list, in a randomised order, in open or close set-format.

5
Auditory Working Memory and Sound Segregation Ability Predict Speech-in-Noise in Adult Cochlear Implant Users

Colak, H.; Guo, X.; Benzaquen, E.; Gurusiddappa, M.; Banerjee, A.; Choi, I.; Sedley, W.; Griffiths, T. D.

2026-06-09 neuroscience 10.64898/2026.06.05.730315 medRxiv
Top 0.1%
19.1%
Show abstract

ObjectivesOutcomes following cochlear implantation vary substantially across adult recipients, and the cognitive and perceptual factors contributing to this variability are not fully understood. This poses a challenge for developing strategies to improve cochlear implant outcomes, as such approaches require a clearer understanding of the mechanisms underlying individual listening difficulties. In this study, we investigated auditory cognitive measures in cochlear implant (CI) users to further elucidate the origins of this variability. DesignThirty-seven adult cochlear implant users completed measures of auditory cognition, comprising auditory working memory (AWM) and sound segregation ability, measured using an auditory figure-ground task (AFG), as well as measures of peripheral temporal and spectral processing, comprising the temporal modulation detection threshold (TMDT) and spectral ripple discrimination threshold (SRDT). Speech perception outcomes were assessed using word-in-noise (WIN) and sentence-in-noise (SIN) tasks. Separate multiple linear regression models evaluated the unique contribution of the auditory cognition measures to WIN and SIN performance, after accounting for the peripheral measures. ResultsBoth regression models explained a substantial proportion of variance in speech-in-noise outcomes (WIN: adjusted R{superscript 2} = 0.55; SIN: adjusted R{superscript 2}=0.57, both p < 0.001). For WIN performance, AFG and AWM were significant predictors. A similar pattern was found for SIN performance, where lower AWM ability and poorer AFG segregation were linked to poorer sentence listening in noise. No significant effects of spectral ripple discrimination or temporal modulation detection were observed in either model, even though both were significantly correlated with WIN performance. ConclusionsThese findings indicate that auditory working memory and sound segregation ability are robust predictors of speech-in-noise outcomes in adult cochlear implant users, across both word- and sentence-level measures. Together, the results may help explain why speech-in-noise outcomes remain highly variable among CI users, even when basic sensory encoding abilities are taken into account. Incorporating measures of auditory working memory and fundamental sound segregation may therefore improve outcome prediction and help in developing more individualised rehabilitation strategies.

6
Learning to localize sounds like a barn owl

Bremen, P.; van der Willigen, R. F.; Van Opstal, A. J.; van Wanrooij, M. M.

2026-02-06 neuroscience 10.64898/2026.02.06.704341 medRxiv
Top 0.1%
18.9%
Show abstract

The brain computes sound location from auditory spatial cues. Humans and barn owls can localize sounds with high accuracy, yet they rely on fundamentally different cue configurations shaped by their ear anatomy and neural circuitry. In humans, symmetrical ears provide interaural time and level differences for horizontal localization, while vertical localization depends primarily on high-frequency, monaural spectral cues generated by the pinnae. Barn owls, by contrast, possess asymmetrical ears and use binaural cues to localize sounds in both azimuth and elevation. Because auditory pathways are assumed to be tuned to the statistics of species-specific cues, it remains unclear whether humans can localize sounds using barn-owl-like spatial information. We addressed this by fitting human listeners with asymmetric ear molds that disrupted normal spectral cues and introduced elevation-dependent interaural level differences, while preserving interaural time differences. Participants wore the molds during daily life and were tested on sound localization using broadband, high-pass, and low-pass noise. Acute exposure to the molds severely degraded elevation localization, while horizontal localization remained largely unaffected. With prolonged exposure, elevation localization improved, but adaptation was limited. Crucially, improvement was strongest for broadband sounds. Because broadband sounds uniquely provide access to both low-frequency interaural time differences and high-frequency interaural level differences, this pattern indicates that listeners learned to use binaural cues to infer sound elevation. These findings demonstrate that the human auditory system can partially adapt to extreme barn-owl-like outer-ear acoustics. Binaural cues can be repurposed to support elevation localization, with effective learning requiring access to complementary spatial cues. Conceptual takeawayGive humans asymmetric barn-owl ears and they can learn to use them - but only partially. The auditory brain is flexible enough to reinterpret new ear shapes, yet is strongly constrained in how far it can relearn a fundamentally different auditory space.

7
Visual contributions to the perception of speech in noise

Alampounti, L. C.; Rosen, S.; Cooper, H.; Bizley, J. K.

2025-09-27 neuroscience 10.1101/2025.09.26.678636 medRxiv
Top 0.1%
18.8%
Show abstract

Investigations of the role of audiovisual integration in speech-in-noise perception have largely focused on the benefits provided by lipreading cues. Nonetheless, audiovisual temporal coherence can offer a complementary advantage in auditory selective attention tasks. We developed an audiovisual speech-in-noise test to assess the benefit of visually conveyed phonetic information and visual contributions to auditory streaming. The test was a video version of the Childrens Coordinate Response Measure with a noun as the second keyword (vCCRMn). The vCCRMn allowed us to measure speech reception thresholds in the presence of two competing talkers under three visual conditions: a full naturalistic video (AV), a video which was interrupted during the target word presentation (Inter), thus, providing no lipreading cues, and a static image of a talker with audio only (A). In each case, the video/image could display either the target talker, or one of the two competing maskers. We assessed speech reception thresholds in each visual condition in 37 young ([&le;] 35 years old) normal-hearing participants. Lipreading ability was independently assessed with the Test of Adult Speechreading (TAS). Results showed that both target-coherent AV and Inter visual conditions offer participants a listening benefit over the static image audio-only condition, with the full AV target-coherent condition providing the most benefit. Lipreading ability correlated with the audiovisual benefit shown in the full AV target-coherent condition, but not the benefit in the Inter target-coherent condition. Together our results are consistent with visual information providing independent benefits to listening, through lip reading and enhanced auditory streaming.

8
Development and clinical application of a consonant confusion task to evaluate hearing aid benefit

Hajicek, J.; Harris, S. E.; Neely, S. T.

2026-04-24 otolaryngology 10.64898/2026.04.23.26351598 medRxiv
Top 0.1%
18.8%
Show abstract

PurposeThis research sought to develop a low-cognitive-load speech-in-noise test based on consonant confusions with the potential for assessing hearing-aid benefit. MethodsVowel-consonant-vowel (VCV) stimuli with added speech-shaped noise were presented as a closed-set consonant identification task. Initially, consonant-confusion matrices were used to select, from a larger set of consonants and vowel contexts, a set of ten consonants and associated signal-to-noise ratios (SNR) that were sensitive to hearing loss. The sensitivity of the qVCV test to hearing loss was validated by comparing predicted pure-tone average (PTA) hearing thresholds with their audiometric PTA. Clinical viability of the qVCV test was assessed by comparisons to the QuickSIN test. Hearing-aid benefit was assessed by comparing test scores in unaided and aided conditions. ResultsThe consonants most sensitive to hearing loss were /b d g t k v z s [esh] n/ in the vowel context /[a]/. A cross-validated prediction of PTA had a mean-absolute error of 5.7 dB. The repeatability of qVCV at 50 trials was equivalent to the QuickSIN average of two lists. Hearing-aid benefit was quantified as a decibel reduction in hearing loss. ConclusionsqVCV and QuickSIN performed similarly when test times are equated. The advantages of qVCV include lower cognitive demand, fewer learning effects, and automated scoring. PTA predicted by qVCV which greatly exceeds audiometric PTA may indicate either cognitive deficits or cochlear neural degeneration. The qVCV quantification of hearing-aid benefit may have clinical value.

9
Auditory deprivation during development alters efferent neural feedback and perception

Mishra, S. K.; Moore, D. R.

2022-11-29 otolaryngology 10.1101/2022.11.24.22282369 medRxiv
Top 0.1%
18.6%
Show abstract

Auditory experience plays a critical role in hearing development. Developmental auditory deprivation due to otitis media, a common childhood disease, produces long-standing changes in the central auditory nervous system, even after the middle ear pathology is resolved. The effects of sound deprivation due to otitis media have been mostly studied in the ascending neural system but remain to be examined in the descending pathway that runs from the auditory cortex to the cochlea via the brainstem. Alterations in the efferent neural system could be important because the descending olivocochlear pathway influences the neural representation of transient sounds in noise in the afferent auditory system and is thought to be involved in auditory relearning following injury. The main objectives of the present study were to (1) investigate whether degraded auditory input due to otitis media during childhood is associated with weakened medial olivocochlear efferent neural responses, even after the resolution of the middle ear pathology, and (2) to examine the involvement of the efferent neural feedback in perceptual masking deficits associated with auditory deprivation due to otitis media. We measured contralateral inhibition of otoacoustic emissions--a biomarker for medial efferent activity--and speech-in-noise recognition in children with a medical history of otitis media (N=76) and age-matched controls (N=99). All children had normal auditory function at the time of experimentation. We found that the inhibitory strength of the medial olivocochlear efferents is weaker in children with a documented history of otitis media relative to controls. In addition, children with otitis media history required a more advantageous signal-to-noise ratio than controls to achieve the same criterion performance level. Importantly, the deficits in perceptual masking were related to efferent inhibition, and these effects could not be attributed to the middle ear or cochlear mechanics. These findings raise the possibility that perceptual masking deficits--a hallmark of impaired (central) auditory processing--resulting from otitis media can arise from the altered brainstem efferent feedback. To date, it was known that degraded auditory experience reorganizes the ascending neural pathways; here, we show that the lack of optimal auditory input to the afferent system during development could have a long-standing impact on the functioning of the descending neural pathways.

10
A Sensory-Cognitive Dissociation in Listeners with Hearing Difficulties: An Exploratory Analysis Linking Tinnitus to Binaural Unmasking Deficits and Speech Complaints to Memory

Bleeck, S.; Hamza, Y.

2025-12-19 otolaryngology 10.64898/2025.12.18.25342552 medRxiv
Top 0.1%
18.5%
Show abstract

BackgroundThe construct of Hidden Hearing Loss (HHL) proposes a link between patient-reported hearing difficulties and underlying neural deficits not captured by the standard audiogram. However, the heterogeneity of this population challenges the utility of HHL as a unitary diagnosis. This study presents an exploratory analysis aimed at deconstructing the HHL symptom complex. MethodsIn 30 participants with a range of hearing abilities and complaints, we measured binaural unmasking using the Binaural Intelligibility Level Difference (BILD). We employed a two-stage analysis. First, a "lumping" analysis tested whether participants could be grouped into a unitary "HHL profile" that predicted a BILD deficit, using both theory-driven classification and data-driven clustering. Second, after this approach failed, a pre-planned exploratory "splitting" analysis used a Linear Mixed-Effects Model (LMM) to investigate whether individual clinical markers (tinnitus, self-reported speech difficulty) were independently associated with the BILD. ResultsThe "lumping" analyses failed to find a significant difference in the BILD between subgroups, questioning the utility of a unitary HHL profile. In contrast, the exploratory "splitting" analysis found a significant interaction between tinnitus and listening condition ({beta} = 1.57, p = 0.009), suggesting that participants with tinnitus exhibited a smaller BILD. The complaint of speech perception difficulty was not significantly associated with a BILD deficit (p = 0.086) but was associated with lower scores on a test of short-term memory (forward digit span, p = 0.046). ConclusionOur findings challenge the value of a unitary HHL profile for predicting this specific binaural deficit. Instead, our exploratory analysis generated a specific, testable hypothesis of a sensory-cognitive dissociation: in our sample, tinnitus was associated with a reduced capacity for binaural unmasking, while the complaint of speech difficulty was associated with poorer short-term memory. These preliminary findings, derived from post-hoc analysis of an underpowered study, require rigorous validation in larger, pre-registered studies.

11
Informational masking vs. crowding - A mid-level trade-off between auditory and visual processing

Zhang, M.; Denison, R. N.; Pelli, D. G.; Le, T. T. C.; Ihlefeld, A.

2021-04-22 neuroscience 10.1101/2021.04.21.440826 medRxiv
Top 0.1%
18.4%
Show abstract

In noisy or cluttered environments, sensory cortical mechanisms help combine auditory or visual features into perceived objects. Knowing that individuals vary greatly in their ability to suppress unwanted sensory information, and knowing that the sizes of auditory and visual cortical regions are correlated, we wondered whether there might be a corresponding relation between an individuals ability to suppress auditory vs. visual interference. In auditory masking, background sound makes spoken words unrecognizable. When masking arises due to interference at central auditory processing stages, beyond the cochlea, it is called informational masking (IM). A strikingly similar phenomenon in vision, called visual crowding, occurs when nearby clutter makes a target object unrecognizable, despite being resolved at the retina. We here compare susceptibilities to auditory IM and visual crowding in the same participants. Surprisingly, across participants, we find a negative correlation (R = -0.7) between IM susceptibility and crowding susceptibility: Participants who have low susceptibility to IM tend to have high susceptibility to crowding, and vice versa. This reveals a mid-level trade-off between auditory and visual processing.

12
Hearing in categories aids speech streaming at the "cocktail party"

Bidelman, G.; Bernard, F.; Skubic, K.

2024-04-05 neuroscience 10.1101/2024.04.03.587795 medRxiv
Top 0.1%
18.4%
Show abstract

Our perceptual system bins elements of the speech signal into categories to make speech perception manageable. Here, we aimed to test whether hearing speech in categories (as opposed to a continuous/gradient fashion) affords yet another benefit to speech recognition: parsing noisy speech at the "cocktail party." We measured speech recognition in a simulated 3D cocktail party environment. We manipulated task difficulty by varying the number of additional maskers presented at other spatial locations in the horizontal soundfield (1-4 talkers) and via forward vs. time-reversed maskers, promoting more and less informational masking (IM), respectively. In separate tasks, we measured isolated phoneme categorization using two-alternative forced choice (2AFC) and visual analog scaling (VAS) tasks designed to promote more/less categorical hearing and thus test putative links between categorization and real-world speech-in-noise skills. We first show that listeners can only monitor up to [~]3 talkers despite up to 5 in the soundscape and streaming is not related to extended high-frequency hearing thresholds (though QuickSIN scores are). We then confirm speech streaming accuracy and speed decline with additional competing talkers and amidst forward compared to reverse maskers with added IM. Dividing listeners into "discrete" vs. "continuous" categorizers based on their VAS labeling (i.e., whether responses were binary or continuous judgments), we then show the degree of IM experienced at the cocktail party is predicted by their degree of categoricity in phoneme labeling; more discrete listeners are less susceptible to IM than their gradient responding peers. Our results establish a link between speech categorization skills and cocktail party processing, with a categorical (rather than gradient) listening strategy benefiting degraded speech perception. These findings imply figure-ground deficits common in many disorders might arise through a surprisingly simple mechanism: a failure to properly bin sounds into categories.

13
Frequency transfer of the ventriloquism aftereffect

Ege, R.; van Opstal, A. J.; van Wanrooij, M. M.

2021-12-23 neuroscience 10.1101/2021.12.22.473801 medRxiv
Top 0.1%
18.3%
Show abstract

Humans localise sounds in the horizontal plane by processing level and timing differences between the ears. This neurocomputational process is continuously and adaptively calibrated using visual input, as seen in the ventriloquism after effect: a shift in sound perception toward a previously seen light. It is unknown from where in the brain this aftereffect originates; adaptation could occur at an early level in the auditory system where neurons are narrowly tuned to frequency, at a later level in the auditory system where localisation cues are extracted, or outside the auditory system at a higher-level spatial map. To investigate this, we examined how the ventriloquism aftereffect generalises across sound frequencies. Participants localised seven narrowband sounds (0.5-8 kHz), targeting different localisation cues. We found that sound localisation accuracy in darkness varied slightly with frequency. When sounds were paired with a visual stimulus that was offset by 10 deg, participants exhibited a pronounced bias toward the light of about [~]63%, corresponding to the well-known ventriloquism effect. The bias was stronger for narrowband compared to broadband sounds. After exposure to a block of these audiovisual stimuli, a ventriloquism aftereffect in the form of a spatial bias of [~]12% was observed across all tested frequencies, largely independent of the frequency of the exposure sound. Together with earlier reports of both frequency-specific and frequency-general recalibration, our results indicate that under conditions of a fixed and consistent audiovisual spatial offset, the ventriloquism aftereffect generalises across sound frequencies, consistent with adaptation at a frequency-independent multisensory spatial stage.

14
Feasibility of efficient smartphone-based threshold and loudness assessments in typical home settings

Xu, C.; Schell-Majoor, L.; Kollmeier, B.

2024-11-19 otolaryngology 10.1101/2024.11.19.24317529 medRxiv
Top 0.1%
18.2%
Show abstract

Reliable hearing assessment at home can improve accessibility and reduce dependence on in-clinic testing. To be viable, home-based procedures must provide accurate results within short measurement times and remain robust to factors such as ambient noise and variable user attention. This study validated two such procedures--a Graded Response Bracketing method for pure-tone threshold estimation and a reinforced adaptive categorical loudness scaling method for loudness-growth assessment--using remote, smartphone-based testing. Fifteen young adults with normal hearing completed the tasks at home and in the laboratory. Ambient noise levels in home environments were also recorded. Test-retest reliability was assessed by repeating the home measurements on a separate occasion. Remote measurements closely matched laboratory results, with mean differences below 1 dB for threshold estimation and below 5 dB for loudness scaling. Test-retest differences obtained at home were small, remaining below 2 dB for threshold estimation and below 1 dB for loudness scaling. These findings demonstrate that smartphone-based pure-tone audiometry and loudness-scaling assessments can achieve high accuracy, efficiency, and reliability when using these procedures, provided that basic acoustic-hygiene conditions (e.g., sufficiently low ambient noise) are maintained.

15
Increased listening effort and decreased speech discrimination at high presentation sound levels in acoustic hearing listeners and cochlear implant users

Huang, C. G.; Field, N. A.; Latorre, M.-E.; Anderson, S.; Goupell, M. J.

2024-09-21 neuroscience 10.1101/2024.09.20.614145 medRxiv
Top 0.1%
18.2%
Show abstract

The sounds we experience in our everyday communication can vary greatly in terms of level and background noise depending on the environment. Paradoxically, increasing the sound intensity may lead to worsened speech understanding, especially in noise. This is known as the "Rollover" phenomenon. There have been limited studies on rollover and how it is experienced differentially across aging groups, for those with and without hearing loss, as well as cochlear implant (CI) users. There is also mounting evidence that listening effort plays an important role in challenging listening conditions and can be directly quantified with objective measures such as pupil dilation. We found that listening effort was modulated by sound level and that rollover occurred primarily in the presence of background noise. The effect on listening effort was exacerbated by age and hearing loss in acoustic listeners, with greatest effect in older listeners with hearing loss, while there was no effect in CI users. The age- and hearing-dependent effects of rollover highlight the potential negative impact of amplification to high sound levels and therefore has implications for effective treatment of age-related hearing loss.

16
Speech-in-Noise Difficulties in Aminoglycoside Ototoxicity Reflects Combined Afferent and Efferent Dysfunction

Motlagh Zadeh, L.; Izhiman, D.; Blankenship, C. M.; Moore, D. R.; Martin, D. K.; Garinis, A.; Feeney, P.; Hunter, L. R.

2026-03-26 otolaryngology 10.64898/2026.03.23.26348719 medRxiv
Top 0.1%
17.3%
Show abstract

Objectives: Patients with Cystic fibrosis (CF) often receive aminoglycosides (AGs) to manage recurrent pulmonary infections, placing them at risk for ototoxicity. Chronic AG use can lead to complex cochlear damage affecting inner and outer hair cells, the stria vascularis, and spiral ganglion neurons. The greatest damage is typically in the basal cochlear region, which encodes high-frequency hearing, with additional involvement of more apical regions. While extended-high-frequency (EHF) hearing loss (EHFHL; 9-16 kHz) is often the earliest sign of AG ototoxicity, speech in noise (SiN) effects are rarely studied. Our overall hypothesis is that SiN perception difficulties in individuals with CF, treated with AGs, are related to combined cochlear and neural damage, primarily in the EHF range but also in the standard frequency (SF; 0.25-8 kHz) range. Three mechanisms that contribute to SiN perception were evaluated in children and young adults: 1) a primary effect of reduced EHF sensitivity, measured by pure-tone audiometry (PTA) and transient-evoked otoacoustic emissions (TEOAEs); 2) a secondary effect of subclinical damage in the SF range, measured by PTA and TEOAEs; and 3) additional neural effects, measured by middle ear muscle reflex (MEMR) threshold (afferent) and growth functions (efferent).Design:A total of 185 participants were enrolled; 101 individuals with CF treated with intravenous AGs and 84 age and sex-matched Controls without hearing concerns or CF. Assessments included EHF and SF PTA; the Bamford-Kowal-Bench (BKB)-SIN test for SiN perception; double-evoked TEOAEs with chirp stimuli from 0.71 to 14.7 kHz; and ipsilateral and contralateral wideband MEMR thresholds and growth functions using broadband stimuli. Results: Reduced sensitivity at EHFs (PTA, TEOAEs) was not associated with impaired SiN perception in the CF group. SF hearing, regardless of EHF status, was the primary predictor of SiN performance in the CF group. Increased MEMR growth was also significantly associated with poorer SiN in the CF group. Conclusions: In CF, impaired SiN perception was primarily predicted by SF hearing impairment, with additional involvement of the efferent auditory pathway through increased MEMR growth. These results build on prior evidence for efferent neural effects due to ototoxic exposures, supporting both sensory (afferent) and neural (efferent) mechanisms that contribute to listening difficulties in CF. Thus, preventive and intervention strategies should consider these combined mechanisms in people with AG ototoxicity to address their SiN problems.

17
The Impact of Instructions on Individual Prioritization Strategies in a Dual-Task Paradigm for Listening Effort

Kestens, K.; Lepla, E.; Vandoorne, F.; Ceuleers, D.; Van Goylen, L.; Keppler, H.

2024-06-26 otolaryngology 10.1101/2024.06.26.24309528 medRxiv
Top 0.1%
15.7%
Show abstract

IntroductionThis study examined the impact of instructions on the prioritization strategy employed by individuals during a listening effort dual-task paradigm. MethodsThe dual-task paradigm consisted of a primary speech understanding task in different listening conditions and a secondary visual memory task, both performed separately (baseline) and simultaneously (dual-task). Twenty-three normal-hearing participants (mean age: 36.8 years; 14 females) were directed to prioritize the primary speech understanding task in the dual-task condition, whereas another twenty-three (matched for age, gender, and education level) received no specific instructions regarding task priority. Both groups performed the dual-task paradigm twice (mean interval: 14.8 days). Patterns of dual-task interference were assessed by plotting the dual-task effect of the primary and secondary task against each other. Fishers exact tests were used to assess whether there was an association between interference patterns and group (non-prioritizing and prioritizing) across all listening conditions and test sessions. ResultsNo statistically significant association was found between the pattern of dual-task interference and the group to which the participants belong for any of the listening conditions and test sessions. Descriptive analysis revealed no consistent strategy use within individuals across listening conditions and test sessions, suggesting a lack of a uniform approach regardless of the given instructions. ConclusionProviding prioritization instructions was insufficient to ensure that an individual will mainly focus on the primary task and consistently adhere to this strategy across listening conditions and test sessions. These results raised reservations about the current usage of dual-task paradigms for listening effort.

18
Performance Analysis of Speech Recognition Models in Automated Scoring of the QuickSIN Test

Hassanpour, A.; Jiang, Y.; Folkeard, P.; Macpherson, E.; Scollie, S. D.; Parsa, V.

2025-07-25 otolaryngology 10.1101/2025.07.25.25332211 medRxiv
Top 0.1%
15.6%
Show abstract

PurposeBest practices in audiology recommend assessing speech understanding in noisy environments, especially for those with communication difficulties. Speech-in-noise (SiN) assessments such as the QuickSIN are used for validating signal processing in hearing aids (HAs) and are linked to HA satisfaction. This project seeks to enhance QuickSIN test efficiency by applying recent advancements in automatic speech recognition (ASR) technologies. MethodTwenty-three adults with sensorineural hearing loss were fitted bilaterally with Unitron Moxi HAs and were administered the QuickSIN test in low and high reverberation environments. Testing was performed with two different HA programs: an omnidirectional program and a fixed directional microphone program. QuickSIN sentences were presented from 0{degrees} azimuth and competing babble from either 0{degrees}, laterally from 90{degrees} or 270{degrees}, or simultaneously from 90{degrees}, 180{degrees}, and 270{degrees} azimuths. Participants verbal responses to QuickSIN stimuli were scored by an audiologist and were recorded in parallel for offline transcription and scoring by ASR models from Amazon, Microsoft, NVIDIA, and Picovoice. The ASR-derived QuickSIN scores were compared to the corresponding audiologist-derived scores. ResultsRepeated Measures ANOVA results revealed that all ASR models overestimated the QuickSIN scores across most test conditions. Bland-Altman analyses showed that the Amazon ASR model had the least bias and the narrowest range for the limits of agreement, in comparison to the manual scoring by an experienced audiologist. ConclusionsSome ASR models, such as Amazon, demonstrated performance comparable to that of an audiologist in automatically scoring QuickSIN tests. However, further refinements are necessary to increase the robustness of the ASR models in scoring low SNR loss test conditions.

19
Age-related changes in acoustic cue use for speech-in-speech perception

Fish, E.; DiNino, M.

2026-06-22 otolaryngology 10.64898/2026.06.17.26355866 medRxiv
Top 0.1%
15.3%
Show abstract

Acoustic cues such as pitch and spatial location allow listeners to attend to a target speaker and ignore competing talkers, aiding speech recognition in background noise. Diminished ability to utilize acoustic cues for speech stream segregation may thus contribute to older adults' challenges hearing in noise. Adults aged 18-74 completed a speech-in-speech identification task with three conditions containing 1) only pitch cues (fundamental frequency), 2) only spatial cues (interaural time differences; ITDs), and 3) both pitch and spatial cues for segregating a target talker from competing talkers. Hearing thresholds at standard and extended high frequencies (EHFs), auditory brainstem responses (ABRs), and digit span scores were acquired to examine the influence of sensory and cognitive factors on use of each acoustic cue for speech-in-speech recognition. Significant differences were observed between cue condition scores indicating that use of the available cue(s) drove performance. ABR metrics were not a significant predictor but digit span scores significantly predicted scores on all three cue conditions. Working memory abilities therefore set a baseline for participants' speech-in-speech recognition regardless of the acoustic content. Hearing thresholds at standard frequencies significantly predicted scores on the Pitch condition. EHF hearing thresholds better predicted Spatial and Both Cue condition performance, suggesting that EHF thresholds represent auditory processing important for coding ITDs. Age group analysis revealed that older adults (aged 40+) performed significantly more poorly on all cue conditions of the speech-in-speech recognition task relative to younger adults. Age-related changes in auditory sensory processing may therefore impair older adults' speech-in-noise perception by reducing their ability to use acoustic cues for segregating target and competing speech.

20
Auditory grouping ability predicts speech-in-noise performance in cochlear implants

Choi, I.; Gander, P. E.; Berger, J. I.; Hong, J.; Colby, S.; McMurray, B.; Griffiths, T. D.

2022-05-31 otolaryngology 10.1101/2022.05.30.22275790 medRxiv
Top 0.1%
14.7%
Show abstract

ObjectivesCochlear implant (CI) users exhibit a large variance in understanding speech in noise (SiN). Past works in CI users found that spectral and temporal resolutions correlate with the SiN ability, but a large portion of variance has been remaining unexplained. Our groups recent work on normal-hearing listeners showed that the ability of grouping temporally coherent tones in a complex auditory scene predicts SiN ability, highlighting a central mechanism of auditory scene analysis that contributes to SiN. The current study examined whether the auditory grouping ability contributes to SiN understanding in CI users as well. Design47 post-lingually deafened CI users performed multiple tasks including sentence-in-noise understanding, spectral ripple discrimination, temporal modulation detection, and stochastic figure-ground task in which listeners detect temporally coherent tone pips in the cloud of many tone pips that rise at random times at random frequencies. Accuracies from the latter three tasks were used as predictor variables while the sentence-in-noise performance was used as the dependent variable in a multiple linear regression analysis. ResultsNo co-linearity was found between any predictor variables. All the three predictors exhibited significant contribution in the multiple linear regression model, indicating that the ability to detect temporal coherence in a complex auditory scene explains a further amount of variance in CI users SiN performance that was not explained by spectral and temporal resolution. ConclusionsThis result indicates that the across-frequency comparison builds an important auditory cognitive mechanism in CI users SiN understanding. Clinically, this result proposes a novel paradigm to reveal a source of SiN difficulty in CI users and a potential rehabilitative strategy.