Multimodal Deep Learning for Low-Resource Settings: A Vector Embedding Alignment Approach for Healthcare Applications
Restrepo, D.; Wu, C.; Cajas, S. A.; Nakayama, L. F.; Celi, L. A. G.; Lopez, D. M.
Show abstract
ObjectiveLarge-scale multi-modal deep learning models and datasets have revolutionized various domains such as healthcare, underscoring the critical role of computational power. However, in resource-constrained regions like Low and Middle-Income Countries (LMICs), GPU and data access is limited, leaving many dependent solely on CPUs. To address this, we advocate leveraging vector embeddings for flexible and efficient computational methodologies, aiming to democratize multimodal deep learning across diverse contexts. Background and SignificanceOur paper investigates the computational efficiency and effectiveness of leveraging vector embeddings, extracted from single-modal foundation models and multi-modal Vision-Language Models (VLM), for multimodal deep learning in low-resource environments, particularly in health-care applications. Additionally, we propose an easy but effective inference-time method to enhance performance by further aligning image-text embeddings. Materials and MethodsBy comparing these approaches with traditional multimodal deep learning methods, we assess their impact on computational efficiency and model performance using accuracy, F1-score, inference time, training time, and memory usage across 3 medical modalities such as BRSET (ophthalmology), HAM10000 (dermatology), and SatelliteBench (public health). ResultsOur findings indicate that embeddings reduce computational demands without compromising the models performance, and show that our embedding alignment method improves the performance of the models in medical tasks. DiscussionThis research contributes to sustainable AI practices by optimizing computational resources in resource-constrained environments. It highlights the potential of embedding-based approaches for efficient multimodal learning. ConclusionVector embeddings democratize multimodal deep learning in LMICs, especially in healthcare. Our study showcases their effectiveness, enhancing AI adaptability in varied use cases.
Matching journals
The top 7 journals account for 50% of the predicted probability mass.
Similar papers in this journal
Similar papers in this journal
- A novel interpretable deep transfer learning combining diverse learnable parameters for improved T2D prediction based on single-cell gene regulatory networks 96%
- An Assistive Computer Vision Tool to Automatically Detect Changes in Fish Behavior In Response to Ambient Odor 95%
- Bridging Auditory Perception and Natural Language Processing with Semantically informed Deep Neural Networks 94%
Similar papers in this journal
- SimSearch: A Human-in-the-Loop Learning Framework for Fast Detection of Regions of Interest in Microscopy Images 94%
- pathCLIP: Detection of Genes and Gene Relations from Biological Pathway Figures through Image-Text Contrastive Learning 94%
- Evaluating Explanations from AI Algorithms for Clinical Decision-Making: A Social Science-based Approach 93%
Similar papers in this journal
- Deep ensemble multitask classification of emergency medical call incidents combining multimodal data improves emergency medical dispatch 94%
- Uncertainty in Deep Learning for EEG under Dataset Shifts 94%
- Building Large-Scale Registries from Unstructured Clinical Notes using a Low-Resource Natural Language Processing Pipeline 93%
Similar papers in this journal
- Accurately Differentiating COVID-19, Other Viral Infection, and Healthy Individuals Using Multimodal Features via Late Fusion Learning 93%
- One LLM is not Enough: Harnessing the Power of Ensemble Learning for Medical Question Answering 93%
- Information retrieval in an infodemic: the case of COVID-19 publications 92%
"Similar papers" are the closest papers from that journal in the model's embedding space. They show what the match is built on, but the ranking comes mostly from a classifier over the whole training set, not from these examples alone.