Egocentric perception of walking environments using an interactive vision-language system
Tan, H.; Mihailidis, A.; Laschowski, B.
Show abstract
Large language models can provide a more detailed contextual understanding of a scene beyond what computer vision alone can provide, which have implications for robotics and embodied intelligence. In this study, we developed a novel multimodal vision-language system for egocentric visual perception, with an initial focus on real-world walking environments. We trained a number of state-of-the-art transformer-based vision-language models that use causal language modelling on our custom dataset of 43,055 image-text pairs for few-shot image captioning. We then designed a new speech synthesis model and a user interface to convert the generated image captions into speech for audio feedback to users. Our system also uniquely allows for feedforward user prompts to personalize the generated image captions. Our system is able to generate detailed captions with an average length of 10 words while achieving a high ROUGE-L score of 43.9% and a low word error rate of 28.1% with an end-to-end processing time of 2.2 seconds. Overall, our new multimodal vision-language system can generate accurate and detailed descriptions of natural scenes, which can be further augmented by user prompts. This innovative feature allows our image captions to be personalized to the individual and immediate needs and preferences of the user, thus optimizing the closed-loop interactions between the human and generative AI models for understanding and navigating of real-world environments.
Matching journals
The top 5 journals account for 50% of the predicted probability mass.
Similar papers in this journal
- DAMM for the detection and tracking of multiple animals within complex social and environmental settings 94%
- An Assistive Computer Vision Tool to Automatically Detect Changes in Fish Behavior In Response to Ambient Odor 94%
- Bridging Auditory Perception and Natural Language Processing with Semantically informed Deep Neural Networks 93%
Similar papers in this journal
- Enhanced Cell Tracking Using A GAN-based Super-Resolution Video-to-Video Time-Lapse Microscopy Generative Model 91%
- Enhanced cell segmentation with limited annotated data using generative adversarial networks 91%
- ARBUR, a machine learning-based analysis system for relating behaviors and ultrasonic vocalizations of rats 91%
"Similar papers" are the closest papers from that journal in the model's embedding space. They show what the match is built on, but the ranking comes mostly from a classifier over the whole training set, not from these examples alone.