English
Related papers

Related papers: Building a Video-and-Language Dataset with Human A…

200 papers

The volumetric representation of human interactions is one of the fundamental domains in the development of immersive media productions and telecommunication applications. Particularly in the context of the rapid advancement of Extended…

Computer Vision and Pattern Recognition · Computer Science 2024-02-15 Fatemeh Ghorbani Lohesara , Davi Rabbouni Freitas , Christine Guillemot , Karen Eguiazarian , Sebastian Knorr

The dense, temporal nature of video presents a profound challenge for automated analysis. Despite the use of powerful Vision-Language Models, prevailing methods for video understanding are limited by the inherent disconnect between…

Computer Vision and Pattern Recognition · Computer Science 2026-03-26 Keliang Li , Yansong Li , Hongze Shen , Mengdi Liu , Hong Chang , Shiguang Shan

Multimodal image-language transformers have achieved impressive results on a variety of tasks that rely on fine-tuning (e.g., visual question answering and image retrieval). We are interested in shedding light on the quality of their…

Computation and Language · Computer Science 2021-06-18 Lisa Anne Hendricks , Aida Nematzadeh

In multimodal assistant, where vision is also one of the input modalities, the identification of user intent becomes a challenging task as visual input can influence the outcome. Current digital assistants take spoken input and try to…

Computer Vision and Pattern Recognition · Computer Science 2021-11-02 Alkesh Patel , Joel Ruben Antony Moniz , Roman Nguyen , Nick Tzou , Hadas Kotek , Vincent Renkens

This work presents MAD (Multimodal Affection Dataset), a multimodal emotion dataset designed for affective computing and neurophysiological modeling. MAD is built upon synchronous collection of diverse physiological signals (EEG, ECG, EOG,…

Signal Processing · Electrical Eng. & Systems 2026-03-09 Shengwei Guo , Yunqing Qiao , Wenzhan Zhang , Bo Liu , Yong Wang , Guobing Sun

Building a socially intelligent agent involves many challenges, one of which is to teach the agent to speak guided by its value like a human. However, value-driven chatbots are still understudied in the area of dialogue systems. Most…

Computation and Language · Computer Science 2022-07-25 Liang Qiu , Yizhou Zhao , Jinchao Li , Pan Lu , Baolin Peng , Jianfeng Gao , Song-Chun Zhu

Video-based visual relation detection tasks, such as video scene graph generation, play important roles in fine-grained video understanding. However, current video visual relation detection datasets have two main limitations that hinder the…

Computer Vision and Pattern Recognition · Computer Science 2024-04-09 Tao Wu , Runyu He , Gangshan Wu , Limin Wang

We contribute a comprehensive dataset to study user attention and purchasing behavior on Search Engine Result Pages (SERPs). Previous work has relied on mouse movements as a low-cost large-scale behavioral proxy but also has relied on…

Human-Computer Interaction · Computer Science 2025-07-14 Kayhan Latifzadeh , Jacek Gwizdka , Luis A. Leiva

Current large-scale video datasets focus on general human activity, but lack depth of coverage on fine-grained activities needed to address physical skill learning. We introduce SportSkills, the first large-scale sports dataset geared…

Computer Vision and Pattern Recognition · Computer Science 2026-03-27 Kumar Ashutosh , Chi Hsuan Wu , Kristen Grauman

Human videos contain rich manipulation priors, but using them for robot learning remains difficult because raw observations entangle scene understanding, human motion, and embodiment-specific action. We introduce MoT-HRA, a hierarchical…

Robotics · Computer Science 2026-05-22 Yifan Xie , YuAn Wang , Guangyu Chen , Jinkun Liu , Yu Sun , Wenbo Ding

Recently, multimodal sentiment analysis has seen remarkable advance and a lot of datasets are proposed for its development. In general, current multimodal sentiment analysis datasets usually follow the traditional system of…

Computation and Language · Computer Science 2021-09-20 Hongxuan Tang , Hao Liu , Xinyan Xiao , Hua Wu

We introduce the task of automatic human action co-occurrence identification, i.e., determine whether two human actions can co-occur in the same interval of time. We create and make publicly available the ACE (Action Co-occurrencE) dataset,…

Computer Vision and Pattern Recognition · Computer Science 2024-06-21 Oana Ignat , Santiago Castro , Weiji Li , Rada Mihalcea

Which common human actions and interactions are recognizable in monocular still images? Which involve objects and/or other people? How many is a person performing at a time? We address these questions by exploring the actions and…

Computer Vision and Pattern Recognition · Computer Science 2015-06-09 Matteo Ruggero Ronchi , Pietro Perona

Understanding human actions is a key problem in computer vision. However, recognizing actions is only the first step of understanding what a person is doing. In this paper, we introduce the problem of predicting why a person has performed…

Computer Vision and Pattern Recognition · Computer Science 2016-12-01 Carl Vondrick , Deniz Oktay , Hamed Pirsiavash , Antonio Torralba

Visual-based human action recognition can be found in various application fields, e.g., surveillance systems, sports analytics, medical assistive technologies, or human-robot interaction frameworks, and it concerns the identification and…

Computer Vision and Pattern Recognition · Computer Science 2024-12-23 Antonios Gasteratos , Stavros N. Moutsis , Konstantinos A. Tsintotas , Yiannis Aloimonos

Humans express feelings or emotions via different channels. Take language as an example, it entails different sentiments under different visual-acoustic contexts. To precisely understand human intentions as well as reduce the…

Artificial Intelligence · Computer Science 2021-11-17 Ting Wu , Junjie Peng , Wenqiang Zhang , Huiran Zhang , Chuanshuai Ma , Yansong Huang

Vision-Language Models (VLMs) have achieved impressive progress in perceiving and describing visual environments. However, their ability to proactively reason and act based solely on visual inputs, without explicit textual prompts, remains…

Computer Vision and Pattern Recognition · Computer Science 2026-01-06 Daoan Zhang , Pai Liu , Xiaofei Zhou , Yuan Ge , Guangchen Lan , Jing Bi , Christopher Brinton , Ehsan Hoque , Jiebo Luo

Understanding comprehensive assembly knowledge from videos is critical for futuristic ultra-intelligent industry. To enable technological breakthrough, we present HA-ViD - the first human assembly video dataset that features representative…

Computer Vision and Pattern Recognition · Computer Science 2023-07-13 Hao Zheng , Regina Lee , Yuqian Lu

In recent years, emotion recognition plays a critical role in applications such as human-computer interaction, mental health monitoring, and sentiment analysis. While datasets for emotion analysis in languages such as English have…

Unlike spoken languages where the use of prosodic features to convey emotion is well studied, indicators of emotion in sign language remain poorly understood, creating communication barriers in critical settings. Sign languages present…

Computer Vision and Pattern Recognition · Computer Science 2025-05-26 Phoebe Chua , Cathy Mengying Fang , Takehiko Ohkawa , Raja Kushalnagar , Suranga Nanayakkara , Pattie Maes