Back

Encoding of pretrained large language models mirrors the genetic architectures of human psychological traits.

Xu, B.; Obradovich, N.; Zheng, W.; Loughnan, R.; Shao, L.; Misaki, M.; Thompson, W. K.; Paulus, M.; Fan, C. C.

2025-03-27 health informatics
10.1101/2025.03.27.25324744 medRxiv
Show abstract

Recent advances in large language models (LLMs) have prompted a frenzy in utilizing them as universal translators for biomedical terms. However, the black box nature of LLMs has forced researchers to rely on artificially designed benchmarks without understanding what exactly LLMs encode. We demonstrate that pretrained LLMs can already explain up to 51% of the genetic correlation between items from a psychometrically-validated neuroticism questionnaire, without any fine-tuning. For psychiatric diagnoses, we found disorder names aligned better with genetic relationships than diagnostic descriptions. Our results indicate the pretrained LLMs have encodings mirroring genetic architectures. These findings highlight LLMs potential for validating phenotypes, refining taxonomies, and integrating textual and genetic data in mental health research.

Matching journals

The top 8 journals account for 50% of the predicted probability mass.

50% of probability mass above

"Similar papers" are the closest papers from that journal in the model's embedding space. They show what the match is built on, but the ranking comes mostly from a classifier over the whole training set, not from these examples alone.