Investigating automatic speech emotion recognition for children with autism spectrum disorder in interactive intervention sessions with the social robot Kaspar
Milling, M.; Bartl-Pokorny, K. D.; Schuller, B. W.
Show abstract
In this contribution, we present the analyses of vocalisation data recorded in the first observation round of the European Commissions Erasmus Plus project "EMBOA, Affective loop in Socially Assistive Robotics as an intervention tool for children with autism". In total, the project partners recorded data in 112 robot-supported intervention sessions for children with autism spectrum disorder. Audio data were recorded using the internal and lapel microphone of the H4n Pro Recorder. To analyse the data, we first utilise a child voice activity detection (VAD) system in order to extract child vocalisations from the raw audio data. For each child, session, and microphone, we provide the total time child vocalisations were detected. Next, we compare the results of two different implementations for valence- and arousal-based speech emotion recognition, thereby processing (1) the child vocalisations detected by the VAD and (2) the total recorded audio material. We provide average valence and arousal values for each session and condition. Finally, we discuss challenges and limitations of child voice detection and audio-based emotion recognition in robot-supported intervention settings.
Matching journals
The top 8 journals account for 50% of the predicted probability mass.
Similar papers in this journal
- Interactive Psychometrics for Autism with the Human Dynamic Clamp: Interpersonal Synchrony from Sensory-motor to Socio-cognitive Domains 91%
- Social Perception and Interaction Database - a novel tool to study social cognitive processes with point-light displays. 90%
- Deep Multimodal Representations and Classification of First-Episode Psychosis via Live Face Processing 89%
Similar papers in this journal
Similar papers in this journal
- Speech-driven Facial Animations Improve Speech-in-Noise Comprehension of Humans 94%
- Electromyographic Correlates of Effortful Listening in the Vestigial Auriculomotor System 91%
- The Effect on Speech-in-Noise Perception of Real Faces and Synthetic Faces Generated with either Deep Neural Networks or the Facial Action Coding System 90%
Similar papers in this journal
"Similar papers" are the closest papers from that journal in the model's embedding space. They show what the match is built on, but the ranking comes mostly from a classifier over the whole training set, not from these examples alone.