Back

A large-scale crowd-sourced annotated acoustic dataset of Indian fauna

Ramesh, V.; Singh, S.; Pop, P.; Choksi, P.; Singh, P.; Khanwilkar, S.; Teotia, S.; Burli, P.; Devarajan, K.; A, A.; A Nakhwa, A.; Abdus Shakur, M.; Baishya, R.; Bhagwat, N.; Biniwale, S.; Bora, C.; C S, S.; Chakraborty, N.; D'Souza, S.; D'Souza, E.; Vaishnav, R. D.; Deshpande, K.; Dhanda, A.; G, A.; Ghosh, A.; Goswami, R.; K N, A.; K P, N.; K Rajaraman, B.; K V, G.; Kannan, V.; Karthick, V.; Kotian, M.; Kumar, H.; Kurian, P.; Madhavan, M.; Meena, K.; Mohammad Maslehuddin, A.; Mourya, P.; Mudke, M.; R J, P.; R S Jha, R.; Ramesh, K.; Sailas, S. S.; Sangwan, T.; Mahesh, S.; Satish, R.; Shankar, A

2026-07-21 ecology
10.64898/2026.07.20.739496 bioRxiv
Show abstract

Global rates of biodiversity loss warrant conservation action and monitoring at large geographic scales. Conservation technologies such as acoustic monitoring in conjunction with deep learning now enable us to monitor wildlife simultaneously across space and time. However, for a significant proportion of biodiversity in tropical regions, we cannot yet rely on automated recognition approaches because we lack acoustic templates to robustly train deep learning algorithms. In this paper, we relied on a novel participatory approach, enlisting researchers, conservation practitioners, and nature enthusiasts to create a unique crowd-sourced, open-access dataset of acoustic annotations across taxonomic groups for biodiversity in India. Our dataset comprises 3311 minutes of strongly labelled data (bounding boxes or annotations for a species vocalization) and 2504 minutes of weakly labelled data (indicating the presence of a species within an audio file but lacking bounding boxes) for 518 species across India, spanning 25 of 36 states and union territories. We present metadata and code for data processing and highlight the strengths of a participatory approach to biodiversity monitoring.

Matching journals

The top 1 journal accounts for 50% of the predicted probability mass.

50% of probability mass above

"Similar papers" are the closest papers from that journal in the model's embedding space. They show what the match is built on, but the ranking comes mostly from a classifier over the whole training set, not from these examples alone.