English
Related papers

Related papers: HAIC: Improving Human Action Understanding and Gen…

200 papers

A major challenge in text-video and text-audio retrieval is the lack of large-scale training data. This is unlike image-captioning, where datasets are in the order of millions of samples. To close this gap we propose a new video mining…

Computer Vision and Pattern Recognition · Computer Science 2022-04-05 Arsha Nagrani , Paul Hongsuck Seo , Bryan Seybold , Anja Hauth , Santiago Manen , Chen Sun , Cordelia Schmid

Human-scene vision-language tasks are increasingly prevalent in diverse social applications, yet recent advancements predominantly rely on models specifically tailored to individual tasks. Emerging research indicates that large…

Artificial Intelligence · Computer Science 2024-11-06 Dawei Dai , Xu Long , Li Yutang , Zhang Yuanhui , Shuyin Xia

While there is overall agreement that future technology for organizing, browsing and searching videos hinges on the development of methods for high-level semantic understanding of video, so far no consensus has been reached on the best way…

Computer Vision and Pattern Recognition · Computer Science 2017-06-20 Du Tran , Maksim Bolonkin , Manohar Paluri , Lorenzo Torresani

In the domain of video surveillance, describing the behavior of each individual within the video is becoming increasingly essential, especially in complex scenarios with multiple individuals present. This is because describing each…

Computer Vision and Pattern Recognition · Computer Science 2023-10-05 Lingru Zhou , Yiqi Gao , Manqing Zhang , Peng Wu , Peng Wang , Yanning Zhang

While multimodal large language models (MLLMs) have made substantial progress in single-image spatial reasoning, multi-image spatial reasoning, which requires integration of information from multiple viewpoints, remains challenging.…

Computer Vision and Pattern Recognition · Computer Science 2026-02-11 Masanari Oi , Koki Maeda , Ryuto Koike , Daisuke Oba , Nakamasa Inoue , Naoaki Okazaki

Human activity understanding is crucial for building automatic intelligent system. With the help of deep learning, activity understanding has made huge progress recently. But some challenges such as imbalanced data distribution, action…

Computer Vision and Pattern Recognition · Computer Science 2019-08-07 Yong-Lu Li , Liang Xu , Xinpeng Liu , Xijie Huang , Yue Xu , Mingyang Chen , Ze Ma , Shiyi Wang , Hao-Shu Fang , Cewu Lu

While Multimodal Large Language Models (MLLMs) excel at single-image understanding, they exhibit significantly degraded performance in multi-image reasoning scenarios. Multi-image reasoning presents fundamental challenges including complex…

Computer Vision and Pattern Recognition · Computer Science 2026-01-13 Jianghao Yin , Qingbin Li , Kun Sun , Cheng Ding , Jie Wang , Qin Chen , Jie Zhou , Nan Wang , Changqing Li , Pei Wu , Jian Xu , Zheming Yang , Liang He

With the growing demand for solutions to real-world video challenges, interest in dense video captioning (DVC) has been on the rise. DVC involves the automatic captioning and localization of untrimmed videos. Several studies highlight the…

Computer Vision and Pattern Recognition · Computer Science 2024-12-20 Minkuk Kim , Hyeon Bae Kim , Jinyoung Moon , Jinwoo Choi , Seong Tae Kim

The training of controllable text-to-video (T2V) models relies heavily on the alignment between videos and captions, yet little existing research connects video caption evaluation with T2V generation assessment. This paper introduces…

Artificial Intelligence · Computer Science 2025-05-20 Xinlong Chen , Yuanxing Zhang , Chongling Rao , Yushuo Guan , Jiaheng Liu , Fuzheng Zhang , Chengru Song , Qiang Liu , Di Zhang , Tieniu Tan

Can humans identify AI-generated (fake) videos and provide grounded reasons? While video generation models have advanced rapidly, a critical dimension -- whether humans can detect deepfake traces within a generated video, i.e.,…

With the rapid advancement of Multimodal Large Language Models (MLLMs), they have demonstrated exceptional capabilities across a variety of vision-language tasks. However, current evaluation benchmarks predominantly focus on objective…

Computation and Language · Computer Science 2025-09-24 Haokun Li , Yazhou Zhang , Jizhi Ding , Qiuchi Li , Peng Zhang

Existing long video retrieval systems are trained and tested in the paragraph-to-video retrieval regime, where every long video is described by a single long paragraph. This neglects the richness and variety of possible valid descriptions…

Computer Vision and Pattern Recognition · Computer Science 2024-12-11 Matthew Gwilliam , Michael Cogswell , Meng Ye , Karan Sikka , Abhinav Shrivastava , Ajay Divakaran

The quality of the data and annotation upper-bounds the quality of a downstream model. While there exist large text corpora and image-text pairs, high-quality video-text data is much harder to collect. First of all, manual labeling is more…

Recent advancements in audio-video joint generation models have demonstrated impressive capabilities in content creation. However, generating high-fidelity human-centric videos in complex, real-world physical scenes remains a significant…

Computer Vision and Pattern Recognition · Computer Science 2026-04-21 Lei Zhu , Xing Cai , Yingjie Chen , Yiheng Li , Binxin Yang , Hao Liu , Jie Chen , Chen Li , Jing LYu

Complex activities in real-world audio unfold over extended durations and exhibit hierarchical structure, yet most prior work focuses on short clips and isolated events. To bridge this gap, we introduce MultiAct, a new dataset and benchmark…

Sound · Computer Science 2026-02-09 Peng Zhang , Qingyu Luo , Philip J. B. Jackson , Wenwu Wang

This paper addresses the gap in predicting turn-taking and backchannel actions in human-machine conversations using multi-modal signals (linguistic, acoustic, and visual). To overcome the limitation of existing datasets, we propose an…

Computation and Language · Computer Science 2025-05-21 Yuxin Lin , Yinglin Zheng , Ming Zeng , Wangzheng Shi

For human action understanding, a popular research direction is to analyze short video clips with unambiguous semantic content, such as jumping and drinking. However, methods for understanding short semantic actions cannot be directly…

Computer Vision and Pattern Recognition · Computer Science 2021-11-23 Kenneth Li , Xiao Sun , Zhirong Wu , Fangyun Wei , Stephen Lin

Large Language Models (LLMs) have shown significant limitations in understanding creative content, as demonstrated by Hessel et al. (2023)'s influential work on the New Yorker Cartoon Caption Contest (NYCCC). Their study exposed a…

Computation and Language · Computer Science 2025-02-28 Kuan Lok Zhou , Jiayi Chen , Siddharth Suresh , Reuben Narad , Timothy T. Rogers , Lalit K Jain , Robert D Nowak , Bob Mankoff , Jifan Zhang

Human behavior understanding is arguably one of the most important mid-level components in artificial intelligence. In order to efficiently make use of data, multi-task learning has been studied in diverse computer vision tasks including…

Computer Vision and Pattern Recognition · Computer Science 2018-02-15 Dong-Jin Kim , Jinsoo Choi , Tae-Hyun Oh , Youngjin Yoon , In So Kweon

Caption quality has emerged as a critical bottleneck in training high-quality text-to-image (T2I) and text-to-video (T2V) generative models. While visual language models (VLMs) are commonly deployed to generate captions from visual data,…

Computer Vision and Pattern Recognition · Computer Science 2026-03-10 Varun Ananth , Baqiao Liu , Haoran Cai