English
Related papers

Related papers: Automatically Detecting Confusion and Conflict Dur…

200 papers

Traditional psychological evaluations rely heavily on human observation and interpretation, which are prone to subjectivity, bias, fatigue, and inconsistency. To address these limitations, this work presents a multimodal emotion recognition…

Human-Computer Interaction · Computer Science 2024-12-25 Kris Kraack

Multimodal fusion is crucial in joint decision-making systems for rendering holistic judgments. Since multimodal data changes in open environments, dynamic fusion has emerged and achieved remarkable progress in numerous applications.…

Computer Vision and Pattern Recognition · Computer Science 2024-11-06 Bing Cao , Yinan Xia , Yi Ding , Changqing Zhang , Qinghua Hu

Multimodal emotion recognition aims to recognize emotions for each utterance of multiple modalities, which has received increasing attention for its application in human-machine interaction. Current graph-based methods fail to…

Computation and Language · Computer Science 2023-11-21 Dongyuan Li , Yusong Wang , Kotaro Funakoshi , Manabu Okumura

In this paper, we present our solutions for emotion recognition in the sub-challenges of Multimodal Emotion Recognition Challenge (MER2024). To mitigate the modal competition issue between audio and text, we adopt an early fusion strategy…

Multimedia · Computer Science 2024-10-01 Mengying Ge , Mingyang Li , Dongkai Tang , Pengbo Li , Kuo Liu , Shuhao Deng , Songbai Pu , Long Liu , Yang Song , Tao Zhang

Vision-language models like CLIP have achieved remarkable progress in cross-modal representation learning, yet suffer from systematic misclassifications among visually and semantically similar categories. We observe that such confusion…

Computer Vision and Pattern Recognition · Computer Science 2026-03-04 Maoyuan Shao , Yutong Gao , Xinyang Huang , Chuang Zhu , Lijuan Sun , Guoshun Nan

Visual In-Context Learning (VICL) has emerged as a powerful paradigm, enabling models to perform novel visual tasks by learning from in-context examples. The dominant "retrieve-then-prompt" approach typically relies on selecting the single…

Computer Vision and Pattern Recognition · Computer Science 2026-01-16 Wenwen Liao , Jianbo Yu , Yuansong Wang , Qingchao Jiang , Xiaofeng Yang

Labeling videos at scale is impractical. Consequently, self-supervised visual representation learning is key for efficient video analysis. Recent success in learning image representations suggests contrastive learning is a promising…

Computer Vision and Pattern Recognition · Computer Science 2021-05-03 Nishant Rai , Ehsan Adeli , Kuan-Hui Lee , Adrien Gaidon , Juan Carlos Niebles

Machine learning models are widely used to support stealth assessment in digital learning environments. Existing approaches typically rely on abstracted gameplay log data, which may overlook subtle behavioral cues linked to learners'…

Machine Learning · Computer Science 2025-07-31 Clemens Witt , Thiemo Leonhardt , Nadine Bergner , Mareen Grillenberger

Various state-of-the-art self-supervised visual representation learning approaches take advantage of data from multiple sensors by aligning the feature representations across views and/or modalities. In this work, we investigate how…

Computer Vision and Pattern Recognition · Computer Science 2022-11-28 Thomas M. Hehn , Julian F. P. Kooij , Dariu M. Gavrila

The development of various sensing technologies is improving measurements of stress and the well-being of individuals. Although progress has been made with single signal modalities like wearables and facial emotion recognition, integrating…

Computer Vision and Pattern Recognition · Computer Science 2023-11-08 Majid Hosseini , Morteza Bodaghi , Ravi Teja Bhupatiraju , Anthony Maida , Raju Gottumukkala

Visual recognition inside the vehicle cabin leads to safer driving and more intuitive human-vehicle interaction but such systems face substantial obstacles as they need to capture different granularities of driver behaviour while dealing…

Computer Vision and Pattern Recognition · Computer Science 2022-04-12 Alina Roitberg , Kunyu Peng , Zdravko Marinov , Constantin Seibold , David Schneider , Rainer Stiefelhagen

Speech emotion recognition is a challenging problem because human convey emotions in subtle and complex ways. For emotion recognition on human speech, one can either extract emotion related features from audio signals or employ speech…

Computation and Language · Computer Science 2020-04-06 Haiyang Xu , Hui Zhang , Kun Han , Yun Wang , Yiping Peng , Xiangang Li

Human communication is multi-modal; e.g., face-to-face interaction involves auditory signals (speech) and visual signals (face movements and hand gestures). Hence, it is essential to exploit multiple modalities when designing machine…

Computer Vision and Pattern Recognition · Computer Science 2024-09-05 Marah Halawa , Florian Blume , Pia Bideau , Martin Maier , Rasha Abdel Rahman , Olaf Hellwich

Predicting agents' future trajectories plays a crucial role in modern AI systems, yet it is challenging due to intricate interactions exhibited in multi-agent systems, especially when it comes to collision avoidance. To address this…

Robotics · Computer Science 2021-03-29 Xu Xie , Chi Zhang , Yixin Zhu , Ying Nian Wu , Song-Chun Zhu

In face-to-face dialogues, the form-meaning relationship of co-speech gestures varies depending on contextual factors such as what the gestures refer to and the individual characteristics of speakers. These factors make co-speech gesture…

Computer Vision and Pattern Recognition · Computer Science 2024-09-18 Esam Ghaleb , Bulat Khaertdinov , Wim Pouw , Marlou Rasenberg , Judith Holler , Aslı Özyürek , Raquel Fernández

The emoticons are symbolic representations that generally accompany the textual content to visually enhance or summarize the true intention of a written message. Although widely utilized in the realm of social media, the core semantics of…

Computer Vision and Pattern Recognition · Computer Science 2024-08-06 Ananya Pandey , Dinesh Kumar Vishwakarma

Can Multimodal Large Language Models (MLLMs) discern confused objects that are visually present but audio-absent? To study this, we introduce a new benchmark, AV-ConfuseBench, which simulates an ``Audio-Visual Confusion'' scene by modifying…

Computer Vision and Pattern Recognition · Computer Science 2025-11-14 Qilang Ye , Wei Zeng , Meng Liu , Jie Zhang , Yupeng Hu , Zitong Yu , Yu Zhou

Large language models (LLMs) enable increasingly capable tutoring-style conversational agents, yet effective tutoring requires sensitivity to learners' affective and cognitive states beyond text alone. Facial expressions provide immediate…

Human-Computer Interaction · Computer Science 2026-04-20 Shuangquan Feng , Laura Fleig , Ruisen Tu , Philip Chi , Edmund Bu , Melinda Ozel , Junhua Ma , Teng Fei , Virginia R. de Sa

The analysis of the collaborative learning process is one of the growing fields of education research, which has many different analytic solutions. In this paper, we provided a new solution to improve automated collaborative learning…

Human-Computer Interaction · Computer Science 2019-04-18 Zhang Guo , Kevin Yu , Rebecca Pearlman , Nassir Navab , Roghayeh Barmaki

Recent years have witnessed astonishing advances in the field of multimodal representation learning, with contrastive learning being the cornerstone for major breakthroughs. Latest works delivered further improvements by incorporating…

Computer Vision and Pattern Recognition · Computer Science 2023-02-22 Chaerin Kong , Nojun Kwak