A1 Refereed original research article in a scientific journal

Modeling human social vision with cinematic stimuli: An integrative approach;




AuthorsSantavirta, Severi; Paranko, Birgitta; Seppälä, Kerttu; Hyönä, Jukka; Nummenmaa, Lauri

PublisherAssociation for Research in Vision and Ophthalmology (ARVO)

Publication year2026

Journal: Journal of Vision

Volume26

Issue5

ISSN1534-7362

DOIhttps://doi.org/10.1167/jov.26.5.9

Publication's open availability at the time of reportingOpen Access

Publication channel's open availability Open Access publication channel

Web address https://doi.org/10.1167/jov.26.5.9

Self-archived copy’s web addresshttps://research.utu.fi/converis/portal/detail/Publication/526838500

Self-archived copy's licenceCC BY NC ND

Self-archived copy's versionPublisher`s PDF

Research data linkhttps://github.com/santavis/social-vision-in-cinema


Abstract

Sociability is central for humans. Visual information ranging from low-level physical features (e.g., luminance) to mid-level semantic information (e.g., face recognition) and high-level social inference (e.g., emotional valence of social interactions) is constantly sampled for navigating the social world. In this study, we utilized large-scale eye tracking during natural vision for mapping how different levels of visual information guide the perception of socially relevant features (social vision) simultaneously. In three experiments, participants (N = 166) watched full-length films and short movie clips with varying social content (total duration: 193 minutes) during eye tracking. To model the association between perceptual features and spatiotemporal gaze parameters (gaze position, gaze synchronization, pupil size and blinking), we extracted 39 stimulus features from the movies, including low-level audiovisual features (e.g., luminance, motion), presence and location of mid-level semantic categories (e.g., faces, objects), and high-level social information (e.g., body movements, pleasantness). Integrative analysis techniques with cross-validation were developed to simultaneously associate the perceptual features with the gaze behavior. Pupil size was modulated by luminance, scene cuts, and emotional arousal while gaze position was most accurately predicted by a combination of the presence of human faces, local motion, and entropy. Faces and eyes were prioritized over other semantic categories, and blinking rate decreased during periods of attentional engagement. Altogether, the results show that human social vision is primarily guided by low-level physical features and mid-level semantic categories, while high-level social features such as emotional arousal primarily modulate pupillary responses.


Downloadable publication

This is an electronic reprint of the original article.
This reprint may differ from the original in pagination and typographic detail. Please cite the original version.




Funding information in the publication
Supported by Turku University Foundation and Alfred Kordelin Foundation grants to SS and by Finnish Governmental Research Funding for Turku University Hospital and for the Western Finland collaborative area to SS and LN and by the European Research Council (ERC-ADV #101141656), Jane and Aatos Erkko Foundation to LN.


Last updated on 29/07/2026 12:58:29 PM