Back

Self-Supervised Behavioral Representations Across the Life Course: A Killifish Case Study

Chang, J.-C.; Komatsu, T. S.; Onami, S.

2026-06-29 animal behavior and cognition
10.64898/2026.06.23.733896 bioRxiv
Show abstract

Self-supervised foundation models of aging are increasingly built from longitudinal data (biobanks, electronic health records, wearables) that is inherently incomplete: no individual is followed across a whole lifetime, and how much of each life is captured varies widely. This raises two linked questions: is it worth modeling an individual's whole life course rather than its current state, and can such a model be built from brief, fragmentary records? No human cohort can settle them, because none offers a complete life to compare against. We turn to the African turquoise killifish (Nothobranchius furzeri), tracked from youth to natural death in publicly released recordings, as a controlled testbed: its complete lifespans provide the full-life reference that human data lacks. On these data we build LifeMAE, a two-stage selfsupervised model: a day encoder that summarizes each day of behavior, then a life-course encoder over the trajectory of those daily summaries. We find that the day encoder alone is already strong: from a single day of behavior it predicts chronological age, separates long- from short-lived individuals (coarsely), and flags nearness to death. Adding the life-course encoder improves on none of the three; each is matched by trivially aggregating the day-level predictions (a smoother for age, an early-life average for lifespan). Near-term mortality seems the exception, where the whole-life model looks far better (AUROC 0.81 to 0.91), but the gain is not behavioral: it reflects where each day falls within the observation window (a cue supplied by the model's encoding of time), and a single-day model given that cue closes the gap at any observation length. For these traits, an individual's place in its life course is legible from a single day: the trajectory stage is unnecessary, and the record it needs is as short as one day, the finest grain our day-level setup resolves. For characterizing a cohort, this favors observing many individuals briefly over tracking a few for long. The result joins a growing body of work in which deep and foundation models, fairly benchmarked, fail to beat deliberately simple baselines. We add a concrete mechanism for the over-optimism: a model's encoding of time can leak the very quantity it predicts, which backwardlooking evaluation mistakes for learned biology, so only evaluation fixed to the moment of prediction is trustworthy.

Matching journals

The top 6 journals account for 50% of the predicted probability mass.

1
Nature
645 papers in training set
Top 0.8%
15.1%
2
PLOS Computational Biology
1863 papers in training set
Top 2%
12.7%
3
Nature Machine Intelligence
70 papers in training set
Top 0.2%
8.9%
4
Communications Psychology
22 papers in training set
Top 0.1%
6.7%
5
eLife
5828 papers in training set
Top 25%
4.8%
6
Nature Aging
60 papers in training set
Top 0.4%
4.3%
50% of probability mass above
7
Nature Methods
385 papers in training set
Top 2%
4.0%
8
Nature Neuroscience
252 papers in training set
Top 2%
4.0%
9
Nature Communications
5641 papers in training set
Top 34%
3.5%
10
Scientific Reports
3612 papers in training set
Top 35%
3.2%
11
PLOS ONE
5266 papers in training set
Top 40%
2.7%
12
Neuron
337 papers in training set
Top 3%
2.6%
13
Proceedings of the National Academy of Sciences
2444 papers in training set
Top 21%
2.4%
14
Nature Human Behaviour
95 papers in training set
Top 0.8%
2.4%
15
Science Advances
1243 papers in training set
Top 21%
1.5%
16
Philosophical Transactions of the Royal Society B: Biological Sciences
72 papers in training set
Top 0.8%
1.4%
17
iScience
1154 papers in training set
Top 22%
1.3%
18
Biology Methods and Protocols
61 papers in training set
Top 1%
1.3%
19
Journal of The Royal Society Interface
235 papers in training set
Top 3%
1.1%
20
Nature Medicine
125 papers in training set
Top 2%
1.1%
21
Patterns
78 papers in training set
Top 2%
1.1%
22
Methods in Ecology and Evolution
176 papers in training set
Top 1%
1.1%
23
BMC Bioinformatics
457 papers in training set
Top 5%
1.0%
24
Science
477 papers in training set
Top 9%
0.8%
25
Neural Networks
35 papers in training set
Top 0.6%
0.8%
26
Briefings in Bioinformatics
354 papers in training set
Top 8%
0.6%
27
Nature Genetics
286 papers in training set
Top 5%
0.6%
28
Bioinformatics
1204 papers in training set
Top 9%
0.6%
29
Frontiers in Artificial Intelligence
20 papers in training set
Top 1.0%
0.6%