Back

Development of a new trauma dataset over 38 years from the Young Finns Study

Saarinen, A.; Asikainen, T.; Lehtimäki, T.; Raitakari, O.; Keltikangas-Järvinen, L.

2026-08-31 psychiatry and clinical psychology
10.64898/2026.08.26.26361417 medRxiv
Show abstract

Background: Previous trauma research includes many limitations, such as the scarcity of pretraumatic health measurements and assessment of traumatic experiences with a broad scope across the lifespan. To respond to these gaps, we aimed to develop a new, prospective, population-based trauma dataset from childhood to middle age. Methods: We used the Young Finns Study that is a population-based, multi-generational, prospective study (n = 3596 for the main generation). It has started in 1980 (baseline assessment) and includes follow-ups in 1983, 1986, 1989, 1992, 1997, 2001, 2007, 2011/2012, and 2018-2020. From the 38-year follow-up and ten measurement points of the YFS, we collected all relevant trauma variables, including both free-format and structured questions that both the participants and their parents responded to. By a data-driven case-to-case analysis, we developed a scale to numerically capture variation in the quality of the experiences. Results: Our final dataset captured a total of 7769 traumatic experiences. We also developed the Traumatic Experience Severity Scale (TESS), including six subscales such as shamefulness, rarity, danger to life or health, effects on everyday life, human-made physical threat, and whether the target person was within or outside one's household. We also preprocessed the dataset to be later easily interleaved with other psychological, cardiovascular, and epigenetic variables of the YFS. Conclusions: We believe this new trauma dataset with thousands of experiences across the lifespan provides new opportunities to multidisciplinary, lifelong trauma research.

Matching journals

The top 6 journals account for 50% of the predicted probability mass.

50% of probability mass above

"Similar papers" are the closest papers from that journal in the model's embedding space. They show what the match is built on, but the ranking comes mostly from a classifier over the whole training set, not from these examples alone.