English
Related papers

Related papers: HAIC: Improving Human Action Understanding and Gen…

200 papers

Existing dense or paragraph video captioning approaches rely on holistic representations of videos, possibly coupled with learned object/action representations, to condition hierarchical language decoders. However, they fundamentally lack…

Computer Vision and Pattern Recognition · Computer Science 2024-01-10 Shih-Han Chou , James J. Little , Leonid Sigal

We present HourVideo, a benchmark dataset for hour-long video-language understanding. Our dataset consists of a novel task suite comprising summarization, perception (recall, tracking), visual reasoning (spatial, temporal, predictive,…

Computer Vision and Pattern Recognition · Computer Science 2024-11-08 Keshigeyan Chandrasegaran , Agrim Gupta , Lea M. Hadzic , Taran Kota , Jimming He , Cristóbal Eyzaguirre , Zane Durante , Manling Li , Jiajun Wu , Li Fei-Fei

Current vision-language multimodal models are well-adapted for general visual understanding tasks. However, they perform inadequately when handling complex visual tasks related to human poses and actions due to the lack of specialized…

Computer Vision and Pattern Recognition · Computer Science 2025-06-03 Dewen Zhang , Wangpeng An , Hayaru Shouno

Training effective AI agents for multi-turn interactions requires high-quality data that captures realistic human-agent dynamics, yet such data is scarce and expensive to collect manually. We introduce APIGen-MT, a two-phase framework that…

Understanding human behavior requires measuring behavioral actions. Due to its complexity, behavior is best mapped onto a rich, semantic structure such as language. Emerging multimodal large language models (MLLMs) are promising candidates,…

Computer Vision and Pattern Recognition · Computer Science 2026-01-27 Haozhe Qi , Shaokai Ye , Alexander Mathis , Mackenzie W. Mathis

This paper introduces QCaption, a novel video captioning and Q&A pipeline that enhances video analytics by fusing three models: key frame extraction, a Large Multimodal Model (LMM) for image-text analysis, and a Large Language Model (LLM)…

Computer Vision and Pattern Recognition · Computer Science 2026-01-13 Jiale Wang , Gee Wah Ng , Lee Onn Mak , Randall Cher , Ng Ding Hei Ryan , Davis Wang

We propose a new task, dataset and model for grounded video caption generation. This task unifies captioning and object grounding in video, where the objects in the caption are grounded in the video via temporally consistent bounding boxes.…

Computer Vision and Pattern Recognition · Computer Science 2024-11-13 Evangelos Kazakos , Cordelia Schmid , Josef Sivic

Generative models have shown significant achievements in audio generation tasks. However, existing models struggle with complex and detailed prompts, leading to potential performance degradation. We hypothesize that this problem stems from…

Learning actions from human demonstration video is promising for intelligent robotic systems. Extracting the exact section and re-observing the extracted video section in detail is important for imitating complex skills because human…

Computer Vision and Pattern Recognition · Computer Science 2021-01-14 Iori Yanokura , Naoki Wake , Kazuhiro Sasabuchi , Katsushi Ikeuchi , Masayuki Inaba

Cognitive science has shown that humans perceive videos in terms of events separated by the state changes of dominant subjects. State changes trigger new events and are one of the most useful among the large amount of redundant information…

Computer Vision and Pattern Recognition · Computer Science 2025-02-04 Yuxuan Wang , Difei Gao , Licheng Yu , Stan Weixian Lei , Matt Feiszli , Mike Zheng Shou

We present the ShareGPT4Video series, aiming to facilitate the video understanding of large video-language models (LVLMs) and the video generation of text-to-video models (T2VMs) via dense and precise captions. The series comprises: 1)…

Computer Vision and Pattern Recognition · Computer Science 2024-06-07 Lin Chen , Xilin Wei , Jinsong Li , Xiaoyi Dong , Pan Zhang , Yuhang Zang , Zehui Chen , Haodong Duan , Bin Lin , Zhenyu Tang , Li Yuan , Yu Qiao , Dahua Lin , Feng Zhao , Jiaqi Wang

Large Multimodal Models (LMMs) have achieved significant progress by extending large language models. Building on this progress, the latest developments in LMMs demonstrate the ability to generate dense pixel-wise segmentation through the…

Computer Vision and Pattern Recognition · Computer Science 2024-09-23 Li Zhou , Xu Yuan , Zenghui Sun , Zikun Zhou , Jingsong Lan

Real-world user-generated videos, especially on platforms like TikTok, often feature rich and intertwined audio visual content. However, existing video captioning benchmarks and models remain predominantly visual centric, overlooking the…

Computer Vision and Pattern Recognition · Computer Science 2025-10-07 Peiran Wu , Yunze Liu , Zhengdong Zhu , Enmin Zhou , Junxiao Shen

Memes have emerged as a powerful form of communication, integrating visual and textual elements to convey humor, satire, and cultural messages. Existing research has focused primarily on aspects such as emotion classification, meme…

Machine Learning · Computer Science 2025-01-24 Shiling Deng , Serge Belongie , Peter Ebert Christensen

Real-world video editing demands not only expert knowledge of cinematic techniques but also multimodal reasoning to select, align, and combine footage into coherent narratives. While recent Large Multimodal Models (LMMs) have shown…

Computer Vision and Pattern Recognition · Computer Science 2026-05-12 Andong Deng , Dawei Du , Zhenfang Chen , Wen Zhong , Fan Chen , Guang Chen , Chia-Wen Kuo , Longyin Wen , Chen Chen , Sijie Zhu

We present a novel approach for discovering human interactions in videos. Activity understanding techniques usually require a large number of labeled examples, which are not available in many practical cases. Here, we focus on recovering…

Computer Vision and Pattern Recognition · Computer Science 2015-02-16 Mehran Khodabandeh , Arash Vahdat , Guang-Tong Zhou , Hossein Hajimirsadeghi , Mehrsan Javan Roshtkhari , Greg Mori , Stephen Se

Human-centric generative models are becoming increasingly popular, giving rise to various innovative tools and applications, such as talking face videos conditioned on text or audio prompts. The core of these capabilities lies in powerful…

Computer Vision and Pattern Recognition · Computer Science 2025-07-15 Donglin Di , He Feng , Wenzhang Sun , Yongjia Ma , Hao Li , Wei Chen , Lei Fan , Tonghua Su , Xun Yang

Humans describe complex scenes with compositionality, using simple text descriptions enriched with links and relationships. While vision-language research has aimed to develop models with compositional understanding capabilities, this is…

Computer Vision and Pattern Recognition · Computer Science 2025-02-28 Yu-Guan Hsieh , Cheng-Yu Hsieh , Shih-Ying Yeh , Louis Béthune , Hadi Pour Ansari , Pavan Kumar Anasosalu Vasu , Chun-Liang Li , Ranjay Krishna , Oncel Tuzel , Marco Cuturi

Generating human-object interactions (HOIs) is critical with the tremendous advances of digital avatars. Existing datasets are typically limited to humans interacting with a single object while neglecting the ubiquitous manipulation of…

Computer Vision and Pattern Recognition · Computer Science 2024-09-12 Xintao Lv , Liang Xu , Yichao Yan , Xin Jin , Congsheng Xu , Shuwen Wu , Yifan Liu , Lincheng Li , Mengxiao Bi , Wenjun Zeng , Xiaokang Yang

Marine videos present significant challenges for video understanding due to the dynamics of marine objects and the surrounding environment, camera motion, and the complexity of underwater scenes. Existing video captioning datasets,…

Computer Vision and Pattern Recognition · Computer Science 2025-09-03 Quang-Trung Truong , Yuk-Kwan Wong , Vo Hoang Kim Tuyen Dang , Rinaldi Gotama , Duc Thanh Nguyen , Sai-Kit Yeung
‹ Prev 1 3 4 5 6 7 10 Next ›