English
Related papers

Related papers: HAIC: Improving Human Action Understanding and Gen…

200 papers

Modeling human cognitive states is essential for advanced artificial intelligence. Existing Large Language Models (LLMs) mainly address isolated tasks such as emotion analysis or stance detection, and fail to capture interactions among…

Computation and Language · Computer Science 2026-04-21 Lin Zhong , Siyu Zhu , Zizhen Yuan , Jinhao Cui , Xinyang Zhao , Lingzhi Wang , Hao Chen , Qing Liao

The increasing variety and quantity of tagged multimedia content on a variety of online platforms offer a unique opportunity to advance the field of human action recognition. In this study, we utilize 283,582 unique, unlabeled TikTok video…

Computer Vision and Pattern Recognition · Computer Science 2024-07-17 Yang Qian , Yinan Sun , Ali Kargarandehkordi , Parnian Azizian , Onur Cezmi Mutlu , Saimourya Surabhi , Pingyi Chen , Zain Jabbar , Dennis Paul Wall , Peter Washington

Human action recognition refers to automatic recognizing human actions from a video clip. In reality, there often exist multiple human actions in a video stream. Such a video stream is often weakly-annotated with a set of relevant human…

Computer Vision and Pattern Recognition · Computer Science 2019-02-07 Qian Wang , Ke Chen

Humans consistently outperform state-of-the-art AI models in action recognition, particularly in challenging real-world conditions involving low resolution, occlusion, and visual clutter. Understanding the sources of this performance gap is…

Computer Vision and Pattern Recognition · Computer Science 2026-03-10 Sadegh Rahmaniboldaji , Filip Rybansky , Quoc C. Vuong , Anya C. Hurlbert , Frank Guerin , Andrew Gilbert

Most existing audio-text retrieval (ATR) methods focus on constructing contrastive pairs between whole audio clips and complete caption sentences, while ignoring fine-grained cross-modal relationships, e.g., short segments and phrases or…

Sound · Computer Science 2025-05-06 Yifei Xin , Yuexian Zou

In this paper, we propose XGC-AVis, a multi-agent framework that enhances the audio-video temporal alignment capabilities of multimodal large models (MLLMs) and improves the efficiency of retrieving key video segments through 4 stages:…

Multimedia · Computer Science 2025-09-30 Yuqin Cao , Xiongkuo Min , Yixuan Gao , Wei Sun , Zicheng Zhang , Jinliang Han , Guangtao Zhai

Manual annotation remains the gold standard for high-quality, dense temporal video datasets, yet it is inherently time-consuming. Vision-language models can aid human annotators and expedite this process. We report on the impact of…

Computer Vision and Pattern Recognition · Computer Science 2026-02-09 Juan Gutiérrez , Victor Gutiérrez , Ángel Mora , Silvia Rodriguez , José Luis Blanco

Multimodal large language models (MLLMs) have made significant strides, yet they face challenges in the medical domain due to limited specialized knowledge. While recent medical MLLMs demonstrate strong performance in lab settings, they…

Computation and Language · Computer Science 2024-10-22 Junda Wang , Yujan Ting , Eric Z. Chen , Hieu Tran , Hong Yu , Weijing Huang , Terrence Chen

The increasing availability of image-text pairs has largely fueled the rapid advancement in vision-language foundation models. However, the vast scale of these datasets inevitably introduces significant variability in data quality, which…

Computer Vision and Pattern Recognition · Computer Science 2024-09-05 Lei Zhang , Fangxun Shu , Tianyang Liu , Sucheng Ren , Hao Jiang , Cihang Xie

While the role of humans is increasingly recognized in machine learning community, representation of and interaction with models in current human-in-the-loop machine learning (HITL-ML) approaches are too low-level and far-removed from…

Computation and Language · Computer Science 2021-09-17 Yiwei Yang , Eser Kandogan , Yunyao Li , Walter S. Lasecki , Prithviraj Sen

Detecting and recognizing human action in videos with crowded scenes is a challenging problem due to the complex environment and diversity events. Prior works always fail to deal with this problem in two aspects: (1) lacking utilizing…

Computer Vision and Pattern Recognition · Computer Science 2020-10-19 Li Yuan , Yichen Zhou , Shuning Chang , Ziyuan Huang , Yunpeng Chen , Xuecheng Nie , Tao Wang , Jiashi Feng , Shuicheng Yan

The rapid advancement of Artificial Intelligence Generated Content (AIGC) technology has propelled audio-driven talking head generation, gaining considerable research attention for practical applications. However, performance evaluation…

Computer Vision and Pattern Recognition · Computer Science 2024-03-12 Weixia Zhang , Chengguang Zhu , Jingnan Gao , Yichao Yan , Guangtao Zhai , Xiaokang Yang

We introduce CameraBench, a large-scale dataset and benchmark designed to assess and improve camera motion understanding. CameraBench consists of ~3,000 diverse internet videos, annotated by experts through a rigorous multi-stage quality…

Computer Vision and Pattern Recognition · Computer Science 2025-09-01 Zhiqiu Lin , Siyuan Cen , Daniel Jiang , Jay Karhade , Hewei Wang , Chancharik Mitra , Tiffany Ling , Yuhan Huang , Sifan Liu , Mingyu Chen , Rushikesh Zawar , Xue Bai , Yilun Du , Chuang Gan , Deva Ramanan

Multimodal large language models (MLLMs) are expected to jointly interpret vision, audio, and language, yet existing video benchmarks rarely assess fine-grained reasoning about human speech. Many tasks remain visually solvable or only…

Computer Vision and Pattern Recognition · Computer Science 2026-04-13 Le Thien Phuc Nguyen , Zhuoran Yu , Samuel Low Yu Hang , Subin An , Jeongik Lee , Yohan Ban , SeungEun Chung , Thanh-Huy Nguyen , JuWan Maeng , Soochahn Lee , Yong Jae Lee

Despite progress in video large language models (Video-LLMs), research on instructional video understanding, crucial for enhancing access to instructional content, remains insufficient. To address this, we introduce InstructionBench, an…

Computer Vision and Pattern Recognition · Computer Science 2025-07-01 Haiwan Wei , Yitian Yuan , Xiaohan Lan , Wei Ke , Lin Ma

Multimodal Large Language Models (MLLMs) excel in vision--language tasks by pre-training solely on coarse-grained concept annotations (e.g., image captions). We hypothesize that integrating fine-grained concept annotations (e.g., object…

Computer Vision and Pattern Recognition · Computer Science 2024-12-10 Xiao Xu , Tianhao Niu , Yuxi Xie , Libo Qin , Wanxiang Che , Min-Yen Kan

When people observe events, they are able to abstract key information and build concise summaries of what is happening. These summaries include contextual and semantic information describing the important high-level details (what, where,…

Computer Vision and Pattern Recognition · Computer Science 2021-05-11 Mathew Monfort , SouYoung Jin , Alexander Liu , David Harwath , Rogerio Feris , James Glass , Aude Oliva

Eye-tracking data reveals valuable insights into users' cognitive states but is difficult to analyze due to its structured, non-linguistic nature. While large language models (LLMs) excel at reasoning over text, they struggle with temporal…

Human-Computer Interaction · Computer Science 2025-07-25 Dongyang Guo , Yasmeen Abdrabou , Enkeleda Thaqi , Enkelejda Kasneci

Automatically describing video, or video captioning, has been widely studied in the multimedia field. This paper proposes a new task of sensor-augmented egocentric-video captioning, a newly constructed dataset for it called MMAC Captions,…

Computer Vision and Pattern Recognition · Computer Science 2021-09-08 Katsuyuki Nakamura , Hiroki Ohashi , Mitsuhiro Okada

High-quality benchmarks are crucial for driving progress in machine learning research. However, despite the growing interest in video generation, there is no comprehensive dataset to evaluate human generation. Humans can perform a wide…

Computer Vision and Pattern Recognition · Computer Science 2025-03-07 Emanuele Bugliarello , Anurag Arnab , Roni Paiss , Pieter-Jan Kindermans , Cordelia Schmid