English
Related papers

Related papers: A Multimodal Dataset of Student Oral Presentations…

200 papers

Wearable sensors, such as smartwatches, have become increasingly prevalent across domains like healthcare, sports, and education, enabling continuous monitoring of physiological and behavioral data. In the context of education, these…

Human-Computer Interaction · Computer Science 2025-12-03 Alvaro Becerra , Pablo Villegas , Ruth Cobos

In this article, we present a novel multimodal feedback framework called MOSAIC-F, an acronym for a data-driven Framework that integrates Multimodal Learning Analytics (MMLA), Observations, Sensors, Artificial Intelligence (AI), and…

Human-Computer Interaction · Computer Science 2025-06-11 Alvaro Becerra , Daniel Andres , Pablo Villegas , Roberto Daza , Ruth Cobos

Lecture slide presentations, a sequence of pages that contain text and figures accompanied by speech, are constructed and presented carefully in order to optimally transfer knowledge to students. Previous studies in multimedia and…

Artificial Intelligence · Computer Science 2022-08-18 Dong Won Lee , Chaitanya Ahuja , Paul Pu Liang , Sanika Natu , Louis-Philippe Morency

This work presents the IMPROVE dataset, a multimodal resource designed to evaluate the effects of mobile phone usage on learners during online education. It includes behavioral, biometric, physiological, and academic performance data…

Human-Computer Interaction · Computer Science 2025-08-29 Roberto Daza , Alvaro Becerra , Ruth Cobos , Julian Fierrez , Aythami Morales

In this paper, a novel dataset is introduced, designed to assess student attention within in-person classroom settings. This dataset encompasses RGB camera data, featuring multiple cameras per student to capture both posture and facial…

Providing timely and actionable feedback on oral presentation slides is challenging in higher education, particularly in large classes where teachers cannot realistically deliver detailed formative feedback before students present. This…

Human-Computer Interaction · Computer Science 2026-05-07 Alvaro Becerra , Diego Gomez , Ruth Cobos

In this study, we introduce YODAS (YouTube-Oriented Dataset for Audio and Speech), a large-scale, multilingual dataset comprising currently over 500k hours of speech data in more than 100 languages, sourced from both labeled and unlabeled…

Computation and Language · Computer Science 2024-06-04 Xinjian Li , Shinnosuke Takamichi , Takaaki Saeki , William Chen , Sayaka Shiota , Shinji Watanabe

Large-scale multimodal models achieve strong results on tasks like Visual Question Answering (VQA), but they are often limited when queries require cultural and visual information, everyday knowledge, particularly in low-resource and…

Studying free-standing conversational groups (FCGs) in unstructured social settings (e.g., cocktail party ) is gratifying due to the wealth of information available at the group (mining social networks) and individual (recognizing native…

Computer Vision and Pattern Recognition · Computer Science 2015-06-24 Xavier Alameda-Pineda , Jacopo Staiano , Ramanathan Subramanian , Ligia Batrinca , Elisa Ricci , Bruno Lepri , Oswald Lanz , Nicu Sebe

Publishing open-source academic video recordings is an emergent and prevalent approach to sharing knowledge online. Such videos carry rich multimodal information including speech, the facial and body movements of the speakers, as well as…

Computation and Language · Computer Science 2024-06-05 Zhe Chen , Heyang Liu , Wenyi Yu , Guangzhi Sun , Hongcheng Liu , Ji Wu , Chao Zhang , Yu Wang , Yanfeng Wang

This study presents high-throughput, real-time multi-agent affective computing framework designed to enhance classroom learning through emotional state monitoring. As large classroom sizes and limited teacher student interaction…

Computer Vision and Pattern Recognition · Computer Science 2026-03-18 Hai Nguyen , Hieu Dao , Hung Nguyen , Nam Vu , Cong Tran

Second language acquisition (SLA) is a complex and dynamic process. Many SLA studies that have attempted to record and analyze this process have typically focused on a single modality (e.g., textual output of learners), covered only a short…

Computation and Language · Computer Science 2024-03-27 Masato Hagiwara , Joshua Tanner

In this article, we present a Web-based System called M2LADS, which supports the integration and visualization of multimodal data recorded in learning sessions in a MOOC in the form of Web-based Dashboards. Based on the edBB platform, the…

Human-Computer Interaction · Computer Science 2023-05-23 Álvaro Becerra , Roberto Daza , Ruth Cobos , Aythami Morales , Mutlu Cukurova , Julian Fierrez

Visual attention is highly fragmented during mobile interactions, but the erratic nature of attention shifts currently limits attentive user interfaces to adapting after the fact, i.e. after shifts have already happened. We instead study…

Human-Computer Interaction · Computer Science 2018-07-26 Julian Steil , Philipp Müller , Yusuke Sugano , Andreas Bulling

This work presents MAD (Multimodal Affection Dataset), a multimodal emotion dataset designed for affective computing and neurophysiological modeling. MAD is built upon synchronous collection of diverse physiological signals (EEG, ECG, EOG,…

Signal Processing · Electrical Eng. & Systems 2026-03-09 Shengwei Guo , Yunqing Qiao , Wenzhan Zhang , Bo Liu , Yong Wang , Guobing Sun

Automatic analysis of teacher and student interactions could be very important to improve the quality of teaching and student engagement. However, despite some recent progress in utilizing multimodal data for teaching and learning…

Computers and Society · Computer Science 2022-12-07 Fangli Xu , Lingfei Wu , KP Thai , Carol Hsu , Wei Wang , Richard Tong

Social interactions play a crucial role in shaping human behavior, relationships, and societies. It encompasses various forms of communication, such as verbal conversation, non-verbal gestures, facial expressions, and body language. In this…

Machine Learning · Computer Science 2026-05-13 Alice Zhang , Callihan Bertley , Dawei Liang , Edison Thomaz

Transforming recorded videos into concise and accurate textual summaries is a growing challenge in multimodal learning. This paper introduces VISTA, a dataset specifically designed for video-to-text summarization in scientific domains.…

Computation and Language · Computer Science 2025-05-27 Dongqi Liu , Chenxi Whitehouse , Xi Yu , Louis Mahon , Rohit Saxena , Zheng Zhao , Yifu Qiu , Mirella Lapata , Vera Demberg

This paper describes a system developed to help University students get more from their online lectures, tutorials, laboratory and other live sessions. We do this by logging their attention levels on their laptops during live Zoom sessions…

Multimedia · Computer Science 2021-01-19 Hyowon Lee , Mingming Liu , Hamza Riaz , Navaneethan Rajasekaren , Michael Scriney , Alan F. Smeaton

Observation of classroom interactions can provide concrete feedback to teachers, but current methods rely on manual annotation, which is resource-intensive and hard to scale. This work explores AI-driven analysis of classroom recordings,…

Computer Vision and Pattern Recognition · Computer Science 2025-12-12 Ivo Bueno , Ruikun Hou , Babette Bühler , Tim Fütterer , James Drimalla , Jonathan Kyle Foster , Peter Youngs , Peter Gerjets , Ulrich Trautwein , Enkelejda Kasneci
‹ Prev 1 2 3 10 Next ›