中文
相关论文

相关论文: HAIC: Improving Human Action Understanding and Gen…

200 篇论文

Automatic video captioning aims for a holistic visual scene understanding. It requires a mechanism for capturing temporal context in video frames and the ability to comprehend the actions and associations of objects in a given timeframe.…

计算机视觉与模式识别 · 计算机科学 2022-12-22 Daniel Lukas Rothenpieler , Shahin Amiriparian

Video captioning is the task of automatically generating a textual description of the actions in a video. Although previous work (e.g. sequence-to-sequence model) has shown promising results in abstracting a coarse description of a short…

计算机视觉与模式识别 · 计算机科学 2018-03-30 Xin Wang , Wenhu Chen , Jiawei Wu , Yuan-Fang Wang , William Yang Wang

Nuanced understanding and the generation of detailed descriptive content for (bimanual) manipulation actions in videos is important for disciplines such as robotics, human-computer interaction, and video content analysis. This study…

计算机视觉与模式识别 · 计算机科学 2023-10-03 Fatemeh Ziaeetabar , Reza Safabakhsh , Saeedeh Momtazi , Minija Tamosiunaite , Florentin Wörgötter

While Multimodal Large Language Models (MLLMs) are adept at answering what is in an image-identifying objects and describing scenes-they often lack the ability to understand how an image feels to a human observer. This gap is most evident…

计算机视觉与模式识别 · 计算机科学 2025-12-01 Yiming Chen , Junlin Han , Tianyi Bai , Shengbang Tong , Filippos Kokkinos , Philip Torr

In this paper, we present an approach for identification of actions within depth action videos. First, we process the video to get motion history images (MHIs) and static history images (SHIs) corresponding to an action video based on the…

计算机视觉与模式识别 · 计算机科学 2019-04-02 Mohammad Farhad Bulbul , Saiful Islam , Hazrat Ali

This paper introduces a novel approach to enhance existing motion captioning methods, which directly map representations of movement to high-level descriptive captions (e.g., ``a person doing jumping jacks"). The existing methods require…

机器学习 · 计算机科学 2025-09-03 Clayton Leite , Yu Xiao

Visual captioning aims to generate textual descriptions given images or videos. Traditionally, image captioning models are trained on human annotated datasets such as Flickr30k and MS-COCO, which are limited in size and diversity. This…

计算机视觉与模式识别 · 计算机科学 2021-03-01 Marimuthu Kalimuthu , Aditya Mogadala , Marius Mosbach , Dietrich Klakow

Videos capture events that typically contain multiple sequential, and simultaneous, actions even in the span of only a few seconds. However, most large-scale datasets built to train models for action recognition in video only provide a…

计算机视觉与模式识别 · 计算机科学 2021-09-29 Mathew Monfort , Bowen Pan , Kandan Ramakrishnan , Alex Andonian , Barry A McNamara , Alex Lascelles , Quanfu Fan , Dan Gutfreund , Rogerio Feris , Aude Oliva

How do two individuals differ when performing the same action? In this work, we introduce Video Action Differencing (VidDiff), the novel task of identifying subtle differences between videos of the same action, which has many applications,…

计算机视觉与模式识别 · 计算机科学 2025-03-12 James Burgess , Xiaohan Wang , Yuhui Zhang , Anita Rau , Alejandro Lozano , Lisa Dunlap , Trevor Darrell , Serena Yeung-Levy

Video generation models have developed rapidly in recent years, where generating natural human motion plays a pivotal role. However, accurately evaluating the quality of generated human motion video remains a significant challenge. Existing…

计算机视觉与模式识别 · 计算机科学 2026-04-29 Bingzi Zhang , Kaisi Guan , Ruihua Song

Recent advancements in Multi-modal Large Language Models (MLLMs) have opened new avenues for applications in Embodied AI. Building on previous work, EgoThink, we introduce VidEgoThink, a comprehensive benchmark for evaluating egocentric…

计算机视觉与模式识别 · 计算机科学 2024-10-16 Sijie Cheng , Kechen Fang , Yangyang Yu , Sicheng Zhou , Bohao Li , Ye Tian , Tingguang Li , Lei Han , Yang Liu

The rapid progress of Multimodal Large Language Models (MLLMs) marks a significant step toward artificial general intelligence, offering great potential for augmenting human capabilities. However, their ability to provide effective…

Large-scale pre-training using egocentric human videos has proven effective for robot learning. However, the models pre-trained on such data can be suboptimal for robot learning due to the significant visual gap between human hands and…

机器人学 · 计算机科学 2026-03-17 Guangrun Li , Yaoxu Lyu , Zhuoyang Liu , Chengkai Hou , Jieyu Zhang , Shanghang Zhang

Image Difference Captioning (IDC) generates natural language descriptions that precisely identify differences between two images, serving as a key benchmark for fine-grained change perception, cross-modal reasoning, and image editing data…

计算机视觉与模式识别 · 计算机科学 2026-05-07 Yuancheng Wei , Haojie Zhang , Linli Yao , Lei Li , Jiali Chen , Tao Huang , Yiting Lu , Duojun Huang , Xin Li , Zhao Zhong

Human-centered dynamic scene understanding plays a pivotal role in enhancing the capability of robotic and autonomous systems, in which Video-based Human-Object Interaction (V-HOI) detection is a crucial task in semantic scene…

计算机视觉与模式识别 · 计算机科学 2024-07-22 Hang Zhang , Wenxiao Zhang , Haoxuan Qu , Jun Liu

Human-Object Interaction (HOI) detection is a challenging computer vision task that requires visual models to address the complex interactive relationship between humans and objects and predict HOI triplets. Despite the challenges posed by…

计算机视觉与模式识别 · 计算机科学 2023-09-19 Yichao Cao , Qingfei Tang , Feng Yang , Xiu Su , Shan You , Xiaobo Lu , Chang Xu

We are committed to learning human skill generators at key-step levels. The generation of skills is a challenging endeavor, but its successful implementation could greatly facilitate human skill learning and provide more experience for…

计算机视觉与模式识别 · 计算机科学 2025-02-13 Yilu Wu , Chenhui Zhu , Shuai Wang , Hanlin Wang , Jing Wang , Zhaoxiang Zhang , Limin Wang

The rapid development of AI-generated content (AIGC) technology has led to the misuse of highly realistic AI-generated images (AIGI) in spreading misinformation, posing a threat to public information security. Although existing AIGI…

计算机视觉与模式识别 · 计算机科学 2025-07-08 Ziyin Zhou , Yunpeng Luo , Yuanchen Wu , Ke Sun , Jiayi Ji , Ke Yan , Shouhong Ding , Xiaoshuai Sun , Yunsheng Wu , Rongrong Ji

Video captioning can be used to assess the video understanding capabilities of Multimodal Large Language Models (MLLMs). However, existing benchmarks and evaluation protocols suffer from crucial issues, such as inadequate or homogeneous…

计算机视觉与模式识别 · 计算机科学 2025-06-16 Linhao Yu , Xinguang Ji , Yahui Liu , Fanheng Kong , Chenxi Sun , Jingyuan Zhang , Hongzhi Zhang , V. W. , Fuzheng Zhang , Deyi Xiong

Recent advancements in visual generation technologies have markedly increased the scale and availability of video datasets, which are crucial for training effective video generation models. However, a significant lack of high-quality,…

计算机视觉与模式识别 · 计算机科学 2025-01-07 Hui Li , Mingwang Xu , Yun Zhan , Shan Mu , Jiaye Li , Kaihui Cheng , Yuxuan Chen , Tan Chen , Mao Ye , Jingdong Wang , Siyu Zhu