Cognitive Effort in Second Language Idiomatic Processing: An Eye-Tracking Dataset
Eye-tracking. Second Language Acquisition. Idiomaticity. Multi-Word Expressions. Language Resources. Psycholinguistics. NLP. Dataset
This paper presents the development, structural organization, and validation of an eyetracking dataset designed to investigate the cognitive processing of idiomatic expressions by second-language (L2) learners. Grounded in the intersection of Psycholinguistics and Natural Language Processing (NLP), specifically within the Modeling Idiomaticity in Human and Artificial Language Processing (MIA) initiative, this resource addresses the critical gap in understanding how non-native speakers navigate the non-compositional nature of idioms. While native speakers typically exhibit hybrid processing strategies, facilitated by direct retrieval of figurative meaning, L2 speakers often rely on compositional parsing, a "literal-first" approach that incurs distinct cognitive costs. This dataset captures these costs through high-fidelity ocular metrics, including fixations and regressions, recorded from Portuguese L1 speakers of English across the full spectrum of Common European Framework of Reference for Languages (CEFR) proficiency levels (A1–C2). We describe the experimental protocol using the Potentially Idiomatic Expressions (PIE) context data, the technical implementation using PsychoPy and the Tobii Pro Spark, and the custom data-processing pipeline developed to interface with the PyGaze library. Preliminary analysis reveals a strong inverse correlation between proficiency and regressive eye movements, validating the dataset’s utility as a benchmark for evaluating both human cognitive models and computational representations of idiomaticity. The dataset is released to support the broader community, particularly tasks associated with the AdMIRe (Advancing Multimodal Idiomaticity Representation) initiative and SemEval shared tasks.