English
Related papers

Related papers: AVCAffe: A Large Scale Audio-Visual Dataset of Cog…

200 papers

Human-robot collaborative assembly systems enhance the efficiency and productivity of the workplace but may increase the workers' cognitive demand. This paper proposes an online and quantitative framework to assess the cognitive workload…

Robotics · Computer Science 2022-07-11 Marta Lagomarsino , Marta Lorenzini , Pietro Balatti , Elena De Momi , Arash Ajoudani

Current roadside perception systems mainly focus on instance-level perception, which fall short in enabling interaction via natural language and reasoning about traffic behaviors in context. To bridge this gap, we introduce RoadSceneVQA, a…

Computer Vision and Pattern Recognition · Computer Science 2025-12-29 Runwei Guan , Rongsheng Hu , Shangshu Chen , Ningyuan Xiao , Xue Xia , Jiayang Liu , Beibei Chen , Ziren Tang , Ningwei Ouyang , Shaofeng Liang , Yuxuan Fan , Wanjie Sun , Yutao Yue

Learning action models from real-world human-centric interaction datasets is important towards building general-purpose intelligent assistants with efficiency. However, most existing datasets only offer specialist interaction category and…

Computer Vision and Pattern Recognition · Computer Science 2025-08-07 Liang Xu , Chengqun Yang , Zili Lin , Fei Xu , Yifan Liu , Congsheng Xu , Yiyi Zhang , Jie Qin , Xingdong Sheng , Yunhui Liu , Xin Jin , Yichao Yan , Wenjun Zeng , Xiaokang Yang

Large vision-language models (LVLMs) excel across diverse tasks involving concrete images from natural scenes. However, their ability to interpret abstract figures, such as geometry shapes and scientific plots, remains limited due to a…

Computer Vision and Pattern Recognition · Computer Science 2024-06-04 Lei Li , Yuqi Wang , Runxin Xu , Peiyi Wang , Xiachong Feng , Lingpeng Kong , Qi Liu

Visual emotion analysis or recognition has gained considerable attention due to the growing interest in understanding how images can convey rich semantics and evoke emotions in human perception. However, visual emotion analysis poses…

Computer Vision and Pattern Recognition · Computer Science 2025-04-10 Rahul Singh Maharjan , Marta Romeo , Angelo Cangelosi

We introduce TempoCave, a novel visualization application for analyzing dynamic brain networks, or connectomes. TempoCave provides a range of functionality to explore metrics related to the activity patterns and modular affiliations of…

Human-Computer Interaction · Computer Science 2020-01-16 Ran Xu , Manu Mathew Thomas , Alex Leow , Olusola Ajilore , Angus G. Forbes

When humans perceive the world, they naturally integrate multiple audio-visual tasks within dynamic, real-world scenes. However, current works such as event localization, parsing, segmentation and question answering are mostly explored…

Computer Vision and Pattern Recognition · Computer Science 2026-03-09 Guangyao Li , Xin Wang , Wenwu Zhu

Speech activity detection (or endpointing) is an important processing step for applications such as speech recognition, language identification and speaker diarization. Both audio- and vision-based approaches have been used for this task in…

The ACII Affective Vocal Bursts Workshop & Competition is focused on understanding multiple affective dimensions of vocal bursts: laughs, gasps, cries, screams, and many other non-linguistic vocalizations central to the expression of…

Audio and Speech Processing · Electrical Eng. & Systems 2022-10-31 Alice Baird , Panagiotis Tzirakis , Jeffrey A. Brooks , Christopher B. Gregory , Björn Schuller , Anton Batliner , Dacher Keltner , Alan Cowen

Speech-based analysis offers a scalable and non-invasive approach for detecting cognitive decline, yet progress has been constrained by the limited availability of clinically validated datasets collected under realistic conditions. We…

Modeling empathy is a complex endeavor that is rooted in interpersonal and experiential dimensions of human interaction, and remains an open problem within AI. Existing empathy datasets fall short in capturing the richness of empathy…

Computation and Language · Computer Science 2024-05-27 Jocelyn Shen , Yubin Kim , Mohit Hulse , Wazeer Zulfikar , Sharifa Alghowinem , Cynthia Breazeal , Hae Won Park

Recent developments in modeling language and vision have been successfully applied to image question answering. It is both crucial and natural to extend this research direction to the video domain for video question answering (VideoQA).…

Computer Vision and Pattern Recognition · Computer Science 2019-06-07 Zhou Yu , Dejing Xu , Jun Yu , Ting Yu , Zhou Zhao , Yueting Zhuang , Dacheng Tao

The affect embedded in video data conveys high-level semantic information about the content and has direct impact on the understanding and perception of reviewers, as well as their emotional responses. Affective Video Content Analysis…

Multimedia · Computer Science 2018-06-04 Ligang Zhang , Jiulong Zhang

The rapid surge of text-to-speech and face-voice reenactment models makes video fabrication easier and highly realistic. To encounter this problem, we require datasets that rich in type of generation methods and perturbation strategy which…

Computer Vision and Pattern Recognition · Computer Science 2025-07-29 Zhixi Cai , Kartik Kuckreja , Shreya Ghosh , Akanksha Chuchra , Muhammad Haris Khan , Usman Tariq , Tom Gedeon , Abhinav Dhall

Achieving realistic, vivid, and human-like synthesized conversational gestures conditioned on multi-modal data is still an unsolved problem due to the lack of available datasets, models and standard evaluation metrics. To address this, we…

Computer Vision and Pattern Recognition · Computer Science 2022-09-21 Haiyang Liu , Zihao Zhu , Naoya Iwamoto , Yichen Peng , Zhengqing Li , You Zhou , Elif Bozkurt , Bo Zheng

In a world where technology is increasingly embedded in our everyday experiences, systems that sense and respond to human emotions are elevating digital interaction. At the intersection of artificial intelligence and human-computer…

Human-Computer Interaction · Computer Science 2025-05-06 Karishma Hegde , Hemadri Jayalath

Audio-visual person recognition (AVPR) has received extensive attention. However, most datasets used for AVPR research so far are collected in constrained environments, and thus cannot reflect the true performance of AVPR systems in…

Computer Vision and Pattern Recognition · Computer Science 2023-07-31 Lantian Li , Xiaolou Li , Haoyu Jiang , Chen Chen , Ruihai Hou , Dong Wang

We introduce Ego4D, a massive-scale egocentric video dataset and benchmark suite. It offers 3,670 hours of daily-life activity video spanning hundreds of scenarios (household, outdoor, workplace, leisure, etc.) captured by 931 unique camera…

In crowdsourcing systems, requesters publish tasks, and interested workers provide answers to get rewards. Worker anonymity motivates participation since it protects their privacy. Anonymity with unlinkability is an enhanced version of…

Cryptography and Security · Computer Science 2024-05-17 Sankarshan Damle , Vlasis Koutsos , Dimitrios Papadopoulos , Dimitris Chatzopoulos , Sujit Gujar

Large experimental Collaborations at the LHC (ALICE, ATLAS, CMS, and LHCb) bring together over 13,000 people from hundreds of institutes over the world. There are many (working) cultures inside these international Collaborations. It is…

Physics and Society · Physics 2024-11-01 Sami Räsänen