An Information-Theoretic Perspective on Multi-LLM Uncertainty Estimation
Kruse, M.; Afshar, M.; Khatwani, S.; Mayampurath, A.; Chen, G.; Gao, Y.
Show abstract
Large language models (LLMs) often behave inconsistently across inputs, indicating uncertainty and motivating the need for its quantification in high-stakes settings. Prior work on calibration and uncertainty quantification often focuses on individual models, overlooking the potential of model diversity. We hypothesize that LLMs make complementary predictions due to differences in training and the Zipfian nature of language, and that aggregating their outputs leads to more reliable uncertainty estimates. To leverage this, we propose MUSE (Multi-LLM Uncertainty via Subset Ensembles), a simple information-theoretic method that uses Jensen-Shannon Divergence to identify and aggregate well-calibrated subsets of LLMs. Experiments on binary prediction tasks demonstrate improved calibration and predictive performance compared to single-model and naive ensemble baselines.
Matching journals
The top 9 journals account for 50% of the predicted probability mass.
Similar papers in this journal
- Neural Collective Matrix Factorization for Integrated Analysis of Heterogeneous Biomedical Data 95%
- ACTIVA: realistic single-cell RNA-seq generation with automatic cell-type identification using introspective variational autoencoders 95%
- Learning Sparse Log-Ratios for High-Throughput Sequencing Data 94%
Similar papers in this journal
- Single-Cell Multi-Modal GAN (scMMGAN) reveals spatial patterns in single-cell data from triple negative breast cancer 94%
- Generating hard-to-obtain information from easy-to-obtain information: applications in drug discovery and clinical inference 93%
- Inferring global-scale temporal latent topics from news reports to predict public health interventions for COVID-19 93%
Similar papers in this journal
- Modular Clinical Decision Support Networks (MoDN)—Updatable, Interpretable, and Portable Predictions for Evolving Clinical Environments 94%
- Uncovering the effects of model initialization on deep model generalization: A study with adult and pediatric chest X-ray images 93%
- Explainable deep learning for disease activity prediction in chronic inflammatory joint diseases 92%
Similar papers in this journal
- Generalising uncertainty improves accuracy and safety of deep learning analytics applied to oncology 94%
- Bridging Auditory Perception and Natural Language Processing with Semantically informed Deep Neural Networks 92%
- Mitigating Machine Learning Bias Between High Income and Low-Middle Income Countries for Enhanced Model Fairness and Generalizability 92%
"Similar papers" are the closest papers from that journal in the model's embedding space. They show what the match is built on, but the ranking comes mostly from a classifier over the whole training set, not from these examples alone.