An Effective Automated Algorithm to Isolate Patient Speech from Conversations with Clinicians
Jaquenoud, T.; Keene, S.; Shlayan, N.; Federman, A.; Pandey, G.
Show abstract
A growing number of algorithms are being developed to automatically identify disorders or disease biomarkers from digitally recorded audio of patient speech. An important step in these analyses is to identify and isolate the patients speech from that of other speakers or noise that are captured in a recording. However, current algorithms, such as diarization, only label the identified speech segments in terms of non-specific speakers, and do not identify the specific speaker of each segment, e.g., clinician and patient. In this paper, we present a novel algorithm that not only performs diarization on clinical audio, but also identifies the patient among the speakers in the recording and returns an audio file containing only the patients speech. Our algorithm first uses pretrained diarization algorithms to separate the input audio into different tracks according to nonspecific speaker labels. Next, in a novel step not conducted in other diarization tools, the algorithm uses the average loudness (quantified as power) of each audio track to identify the patient, and return the audio track containing only their speech. Using a practical expert-based evaluation methodology and a large dataset of clinical audio recordings, we found that the best implementation of our algorithm achieved near-perfect accuracy on two validation sets. Thus, our algorithm can be used for effectively identifying and isolating patient speech, which can be used in downstream expert and/or data-driven analyses.
Matching journals
The top 3 journals account for 50% of the predicted probability mass.
Similar papers in this journal
- Time-adaptive Unsupervised Auditory Attention Decoding Using EEG-based Stimulus Reconstruction 95%
- Off-body Sleep Analysis for Predicting Adverse Behavior in Individuals with Autism Spectrum Disorder 91%
- Deep Sentiment Classification and Topic Discovery on Novel Coronavirus or COVID-19 Online Discussions: NLP Using LSTM Recurrent Neural Network Approach 91%
Similar papers in this journal
Similar papers in this journal
- A recurrent neural network and parallel hidden Markov model algorithm to segment and detect heart murmurs in phonocardiograms 93%
- From theoretical models to practical deployment: A perspective and case study of opportunities and challenges in AI-driven healthcare research for low-income settings 93%
- Multiple Instance Learning Framework can Facilitate Explainability in Murmur Detection 93%
Similar papers in this journal
- Comparison of Two-Talker Attention Decoding from EEG with Nonlinear Neural Networks and Linear Methods 94%
- Deep Learning Restores Speech Intelligibility in Multi-Talker Interference for Cochlear Implant Users 93%
- Bridging Auditory Perception and Natural Language Processing with Semantically informed Deep Neural Networks 93%
Similar papers in this journal
- Linear versus deep learning methods for noisy speech separation for EEG-informed attention decoding 95%
- 'Are you even listening?' - EEG-based decoding of absolute auditory attention to natural speech 95%
- Speech decoding from a small set of spatially segregated minimally invasive intracranial EEG electrodes with a compact and interpretable neural network 94%
"Similar papers" are the closest papers from that journal in the model's embedding space. They show what the match is built on, but the ranking comes mostly from a classifier over the whole training set, not from these examples alone.