Back

Bantu on the brain: Community-informed adaptations of the Macarthur-Bates Communication Development Index to contemporary Zambian languages using item response theory

Latoya, M. G.; Rockers, P. C.; Buumba, C.; Shimaingwa, F.; Thea, D. M.; Zulu, E. M.; Herlihy, J. M.

2025-06-05 pediatrics
10.1101/2025.06.03.25328935 medRxiv
Show abstract

The MacArthur-Bates Communication Development Inventories (MB-CDI) are a widely used set of tools to assess language acquisition in early childhood. Although it has been adapted in 120 languages, there is not a linguistic nor culturally congruent tool for any of the 72 Zambian Bantu languages. This manuscript describes the process to adapt the MB-CDI for two of the most spoken languages in Zambia: Nyanja and Bemba. Using mixed-methods with caregivers of children aged 16 to 30 months in Lusaka, Zambia, we have constructed two Bantu-language adaptations of the MD-CDI: Words and Sentences short form. Two focus group discussions of caregivers were conducted, one in Nyanja (n=10) and one in Bemba (n=10). The goal of the focus groups was to generate a list of commonly heard words in early childhood. Facilitators then administered individual assessment surveys to Bemba (n=77) and Nyanja (n=109) caregivers to determine which words from the generated lists their child (aged 16-30 months) had acquired in productive speech. We then fitted a 2-parameter item response theory model to our data to reflect 100 words, characterized by a range of difficulty and highest ability to effectively distinguish between different levels of language proficiency. Our adapted MB-CDI for Bemba and Nyanja is a language acquisition assessment tool for ages 16-30 months to measure language development in the two most spoken Bantu languages in Zambia.

Matching journals

The top 4 journals account for 50% of the predicted probability mass.

50% of probability mass above

"Similar papers" are the closest papers from that journal in the model's embedding space. They show what the match is built on, but the ranking comes mostly from a classifier over the whole training set, not from these examples alone.