Natural-language retrieval with multimodal embeddings identifies candidate developmental behaviors in caregiver-child recordings
Mwangi, B.; Wu, M.-J.; Mansour, R.; Anzueto, G.; Pagan, A. F.
Show abstract
Background Naturalistic audiovisual recordings of caregiver-child interactions contain rich developmental signals. However, extracting interpretable clinical measures requires resource-intensive manual coding. To address this bottleneck, we evaluated natural-language queries for retrieving specific behavioral moments from these recordings, applying multimodal embeddings as an automated evidence-selection layer. Methods We compared three embedding models (Jina Embeddings v5 Omni, LanguageBind, and Wave7B) for natural-language retrieval directly from audio and video streams, bypassing transcript text. We assessed performance across 27 behavioral targets in 277 caregiver-child recordings (14, 24, and 36 months of age) from the Early Head Start Talkbank corpus, yielding 7,479 recording-target queries. Results Jina Embeddings v5 Omni achieved the highest top-10 retrieval success (text-to-audio 38.3%; text-to-video 36.4%), ahead of LanguageBind (37.0%; 34.5%) and Wave7B (36.1%; 35.0%). Across models, retrieval was substantially more successful for common targets than for rare vocal and gestural behaviors, such as pointing and babbling. By analyzing the spoken words within the retrieved audio clips, we found that Jina accurately ranked the children by their relative vocabulary size at each age (Spearman = 0.68, 0.82, and 0.90 at 14, 24, and 36 months). However, the model severely underestimated the total number of unique words each child used throughout the full session. Conclusion Multimodal embeddings can successfully pinpoint important developmental behaviors and speech patterns within lengthy caregiver-child recordings. However, these systems still struggle to locate rare events. Additionally, while they can accurately rank children by relative vocabulary size, they fail to measure a child's complete vocabulary. We conclude that these models are currently best suited for automated evidence-selection to prioritize relevant segments for expert interpretation rather than acting as an independent replacement for manual behavioral coding or language assessment. Improving the detection of infrequent behaviors and validating these models across external datasets are essential next steps before real-world clinical deployment.
Matching journals
The top 9 journals account for 50% of the predicted probability mass.
Similar papers in this journal
- Identification of social engagement indicators associated with autism spectrum disorder using a game-based mobile application 90%
- Structured Codes and Free-Text Notes: Measuring Information Complementarity in Electronic Health Records 88%
- One LLM is not Enough: Harnessing the Power of Ensemble Learning for Medical Question Answering 87%
Similar papers in this journal
Similar papers in this journal
- Evaluation of a Large Language Model to Identify Confidential Content in Adolescent Encounter Notes 90%
- The lasting effects of the pandemic: A time series analysis of first-time speech delays in kids under 5 years of age 89%
- Impacts of school closures on physical and mental health of children and young people: a systematic review 81%
Similar papers in this journal
- Annotation-preserving machine translation of English corpora to validate Dutch clinical concept extraction tools 90%
- Generative Large Language Models in Electronic Health Records for Patient Care Since 2023: A Systematic Review 90%
- LCD Benchmark: Long Clinical Document Benchmark on Mortality Prediction for Language Models 90%
Similar papers in this journal
"Similar papers" are the closest papers from that journal in the model's embedding space. They show what the match is built on, but the ranking comes mostly from a classifier over the whole training set, not from these examples alone.