An autonomous AI agent for universal behavior analysis
Aljovic, A.; Lin, Z.; Wang, W.; Zhang, X.; Marin-Llobet, A.; Liang, N.; Canales, B.; Lee, J.; Baek, J.; Liu, R.; Li, C.; Li, N.; Liu, J.
Show abstract
Behavior analysis across species represents a fundamental challenge in neuroscience, psychology, and ethology, typically requiring extensive expert knowledge and labor-intensive processes that limit research scalability and accessibility. We introduce BehaveAgent, an autonomous multimodal AI agent designed to automate behavior analysis from video input without retraining or manual intervention. Unlike conventional methods that require manual behavior annotation, video segmentation, task-specific model training, BehaveAgent leverages the reasoning capabilities of multimodal large language models (LLM) to generalize across novel behavioral domains without need for additional training. It integrates LLMs, vision-language models (VLMs), and large-scale visual grounding modules, orchestrated through a multimodal context memory and goal-directed attention mechanism, to enable robust zero-shot visual reasoning across species and experimental paradigms, including plants, insects, rodents, primates, and humans. Upon receiving a video input, BehaveAgent autonomously identifies the correct analysis strategy and performs end-to-end behavior analysis and interpretation without human supervision. Leveraging vision-language representations, it performs general-purpose tracking, pose estimation and segmentation. We demonstrate BehaveAgents universal applicability to autonomously (1) identify the behavioral paradigm and develop an action plan specialized for the identified paradigm, (2) identify relevant subjects and objects, (3) track those features, (4) identify behavioral sequences with explicit reasoning, (5) generate and execute code for targeted analysis and (6) generate comprehensive research reports that integrate behavioral findings with relevant scientific literature. Through interpretable agentic reasoning, BehaveAgent makes its internal decision-making process transparent, clarifying why particular features are tracked or behaviors inferred. By reducing the time and expertise required for behavior analysis, BehaveAgent introduces a scalable, generalizable, and explainable paradigm for advancing biological and behavioral research.
Matching journals
The top 4 journals account for 50% of the predicted probability mass.
Similar papers in this journal
- Large-scale capture of hidden fluorescent labels for training generalizable markerless motion capture models 95%
- Automatic mapping of multiplexed social receptive fields by deep learning and GPU-accelerated 3D videography 94%
- TeraVR Empowers Precise Reconstruction of Complete 3-D Neuronal Morphology in the Whole Brain 94%
Similar papers in this journal
- Pancreatlas: applying an adaptable framework to map the human pancreas in health and disease 91%
- Unraveling the complexity of rat object vision requires a full convolutional network - and beyond 91%
- CellContrast: Reconstructing Spatial Relationships in Single-Cell RNA Sequencing Data via Deep Contrastive Learning 91%
Similar papers in this journal
- pyControl: Open source, Python based, hardware and software for controlling behavioural neuroscience experiments. 95%
- Eye movements reveal spatiotemporal dynamics of visually-informed planning in navigation 93%
- Deep Neural Networks to Register and Annotate Cells in Moving and Deforming Nervous Systems 93%
"Similar papers" are the closest papers from that journal in the model's embedding space. They show what the match is built on, but the ranking comes mostly from a classifier over the whole training set, not from these examples alone.