Back

Probing Hidden States for Calibrated, Alignment-Resistant Predictions in LLMs

Berkowitz, J. S.; Kivelson, S.; Srinivasan, A. S.; Gisladottir, U.; Tsang, K.; Acitores Cortina, J. M.; Kuchi, A.; Patock, J. R.; Czarny, R.; Tatonetti, N. P.

2025-09-27 health informatics
10.1101/2025.09.17.25336018 medRxiv
Show abstract

Scientific applications of large language models (LLMs) demand reliable, well-calibrated predictions, but standard generative approaches often fail to fully access relevant knowledge contained in their internal representations. As a result, models appear less capable than they are, with useful information remaining latent. We present PING (Probing INternal states of Generative models), an open-source framework that trains lightweight probes on frozen, HuggingFace-compatible transformers to deliver structured predictions with minimal compute overhead. Across diverse models and benchmarks including MMLU for broad coverage and MedMCQA for clinical focus, PING matches or exceeds generative accuracy while reducing Expected Calibration Error by up to 96%. Strikingly, on an LLM that has been explicitly safety-tuned to withhold medical information, PING recovered 87% of lost MedMCQA performance while generative accuracy is zero, showing this information still exists in the models latent space. The accompanying pingkit package makes these methods easy to deploy and is available through PyPI.

Matching journals

The top 6 journals account for 50% of the predicted probability mass.

50% of probability mass above

"Similar papers" are the closest papers from that journal in the model's embedding space. They show what the match is built on, but the ranking comes mostly from a classifier over the whole training set, not from these examples alone.