A4 Refereed article in a conference publication

Measuring Social Integration Through Participation: Categorizing Organizations and Leisure Activities in the Displaced Karelians Interview Archive using LLMs;




AuthorsLaato, Joonatan; Schroderus, Veera; Kanerva, Jenna; Kauppi, Jenni; Lummaa, Virpi; Ginter, Filip

EditorsAlves, Diego; Bizzoni, Yuri; Degaetano-Ortlieb, Stefania; Kazantseva, Anna; Pagel, Janis; Szpakowicz, Stan

Conference nameSIGHUM Workshop on Computational Linguistics for Cultural Heritage, Social Sciences, Humanities and Literature

PublisherAssociation for Computational Linguistics

Publication year2026

Book title Proceedings of the 10th Joint SIGHUM Workshop on Computational Linguistics for Cultural Heritage, Social Sciences, Humanities and Literature 2026

First page 111

Last page127

ISBN979-8-89176-373-9

DOIhttps://doi.org/10.18653/v1/2026.latechclfl-1.11

Publication's open availability at the time of reportingOpen Access

Publication channel's open availability Open Access publication channel

Web address https://doi.org/10.18653/v1/2026.latechclfl-1.11

Self-archived copy’s web addresshttps://research.utu.fi/converis/portal/detail/Publication/527004670

Self-archived copy's licenceCC BY

Self-archived copy's versionPublisher`s PDF


Abstract
Digitized historical archives make it possible to study everyday social life on a large scale, but the information extracted directly from text often does not directly allow one to answer the research questions posed by historians or sociologists in a quantitative manner. We address this problem in a large collection of Finnish World War II Karelian evacuee family interviews. Prior work extracted more than 350K mentions of leisure time activities and organizational memberships from these interviews, yielding 71K unique activity and organization names—far too many to analyze directly. We develop a categorization framework that captures key aspects of participation (the kind of activity/organization, how social it typically is, how regularly it happens, and how physically demanding it is). We annotate a gold-standard set to allow for a reliable evaluation, and then test whether large language models can apply the same schema at scale. Using a simple voting approach across multiple model runs, we find that an open-weight LLM can closely match expert judgments. Finally, we apply the method to label the 350K entities, producing a structured resource for downstream studies of social integration and related outcomes.

Downloadable publication

This is an electronic reprint of the original article.
This reprint may differ from the original in pagination and typographic detail. Please cite the original version.





Last updated on 18/08/2026 04:25:23 PM