English
Related papers

Related papers: M$^3$AV: A Multimodal, Multigenre, and Multipurpos…

200 papers

The natural association between visual observations and their corresponding sound provides powerful self-supervisory signals for learning video representations, which makes the ever-growing amount of online videos an attractive source of…

Computer Vision and Pattern Recognition · Computer Science 2021-08-18 Sangho Lee , Jiwan Chung , Youngjae Yu , Gunhee Kim , Thomas Breuel , Gal Chechik , Yale Song

As audio-visual systems increasingly bring immersive and interactive capabilities into our work and leisure activities, so the need for naturalistic test material grows. New volumetric datasets have captured high-quality 3D video, but…

Multimedia · Computer Science 2021-05-04 Hanne Stenzel , Davide Berghi , Marco Volino , Philip J. B. Jackson

The technical complexity of research papers often limits their reach, necessitating more accessible formats like scientific videos to disseminate key insights through engaging narration. However, existing automated methods primarily focus…

Artificial Intelligence · Computer Science 2026-04-22 Xiao Liang , Bangxin Li , Zixuan Chen , Hanyue Zheng , Zhi Ma , Di Wang , Cong Tian , Quan Wang

The scarcity of high-quality, multimodal training data severely hinders the creation of lifelike avatar animations for conversational AI in virtual environments. Existing datasets often lack the intricate synchronization between speech,…

Artificial Intelligence · Computer Science 2024-10-23 Saif Punjwani , Larry Heck

We present PresentAgent, a multimodal agent that transforms long-form documents into narrated presentation videos. While existing approaches are limited to generating static slides or text summaries, our method advances beyond these…

Computer Vision and Pattern Recognition · Computer Science 2025-07-08 Jingwei Shi , Zeyu Zhang , Biao Wu , Yanjie Liang , Meng Fang , Ling Chen , Yang Zhao

Audio-visual speech recognition (AVSR) is a multimodal extension of automatic speech recognition (ASR), using video as a complement to audio. In AVSR, considerable efforts have been directed at datasets for facial features such as…

Computer Vision and Pattern Recognition · Computer Science 2024-07-03 Hao Wang , Shuhei Kurita , Shuichiro Shimizu , Daisuke Kawahara

Instruction tuning has significantly advanced large language models (LLMs) such as ChatGPT, enabling them to align with human instructions across diverse tasks. However, progress in open vision-language models (VLMs) has been limited due to…

Computer Vision and Pattern Recognition · Computer Science 2023-06-09 Lei Li , Yuwei Yin , Shicheng Li , Liang Chen , Peiyi Wang , Shuhuai Ren , Mukai Li , Yazheng Yang , Jingjing Xu , Xu Sun , Lingpeng Kong , Qi Liu

Multi-Modal automatic speech recognition (ASR) techniques aim to leverage additional modalities to improve the performance of speech recognition systems. While existing approaches primarily focus on video or contextual information, the…

Sound · Computer Science 2023-12-27 Haoxu Wang , Fan Yu , Xian Shi , Yuezhang Wang , Shiliang Zhang , Ming Li

We present a novel human annotated dataset for evaluating the ability for visual-language models to generate both short and long descriptions for real-world video clips, termed DeVAn (Dense Video Annotation). The dataset contains 8.5K…

Computer Vision and Pattern Recognition · Computer Science 2024-08-12 Tingkai Liu , Yunzhe Tao , Haogeng Liu , Qihang Fan , Ding Zhou , Huaibo Huang , Ran He , Hongxia Yang

Video captioning which automatically translates video clips into natural language sentences is a very important task in computer vision. By virtue of recent deep learning technologies, e.g., convolutional neural networks (CNNs) and…

Computer Vision and Pattern Recognition · Computer Science 2016-11-18 Junbo Wang , Wei Wang , Yan Huang , Liang Wang , Tieniu Tan

Transforming recorded videos into concise and accurate textual summaries is a growing challenge in multimodal learning. This paper introduces VISTA, a dataset specifically designed for video-to-text summarization in scientific domains.…

Computation and Language · Computer Science 2025-05-27 Dongqi Liu , Chenxi Whitehouse , Xi Yu , Louis Mahon , Rohit Saxena , Zheng Zhao , Yifu Qiu , Mirella Lapata , Vera Demberg

Many everyday tasks, ranging from appliance repair and cooking to car maintenance, require expert knowledge, particularly for complex, multi-step procedures. Despite growing interest in AI agents for augmented reality (AR) assistance,…

Computer Vision and Pattern Recognition · Computer Science 2026-02-03 Lavisha Aggarwal , Vikas Bahirwani , Andrea Colaco

Academic presentation videos have become an essential medium for research communication, yet producing them remains highly labor-intensive, often requiring hours of slide design, recording, and editing for a short 2 to 10 minutes video.…

Computer Vision and Pattern Recognition · Computer Science 2025-10-10 Zeyu Zhu , Kevin Qinghong Lin , Mike Zheng Shou

This paper introduces a new video-and-language dataset with human actions for multimodal logical inference, which focuses on intentional and aspectual expressions that describe dynamic human actions. The dataset consists of 200 videos,…

Computer Vision and Pattern Recognition · Computer Science 2021-06-29 Riko Suzuki , Hitomi Yanaka , Koji Mineshima , Daisuke Bekki

Current methods for learning visually grounded language from videos often rely on text annotation, such as human generated captions or machine generated automatic speech recognition (ASR) transcripts. In this work, we introduce the…

Learning text-video embeddings usually requires a dataset of video clips with manually provided captions. However, such datasets are expensive and time consuming to create and therefore difficult to obtain on a large scale. In this work, we…

Computer Vision and Pattern Recognition · Computer Science 2019-08-01 Antoine Miech , Dimitri Zhukov , Jean-Baptiste Alayrac , Makarand Tapaswi , Ivan Laptev , Josef Sivic

The rapid advancement of AI-generated multimodal video-audio content has raised significant concerns regarding information security and content authenticity. Existing synthetic video datasets predominantly focus on the visual modality…

Computer Vision and Pattern Recognition · Computer Science 2026-04-21 Mengxue Hu , Yunfeng Diao , Changtao Miao , Zhiqing Guo , Jianshu Li , Zhe Li , Joey Tianyi Zhou

This paper presents MAST, a new model for Multimodal Abstractive Text Summarization that utilizes information from all three modalities -- text, audio and video -- in a multimodal video. Prior work on multimodal abstractive text…

Computation and Language · Computer Science 2020-10-19 Aman Khullar , Udit Arora

With the rapid adoption of multimodal large language models (MLLMs) across diverse applications, there is a pressing need for task-centered, high-quality training data. A key limitation of current training datasets is their reliance on…

Computer Vision and Pattern Recognition · Computer Science 2025-11-18 Xiaoyu Lin , Aniket Ghorpade , Hansheng Zhu , Justin Qiu , Dea Rrozhani , Monica Lama , Mick Yang , Zixuan Bian , Ruohan Ren , Alan B. Hong , Jiatao Gu , Chris Callison-Burch

This paper introduces VideoMind, a video-centric omni-modal dataset designed for deep video content cognition and enhanced multi-modal feature representation. The dataset comprises 103K video samples (3K reserved for testing), each paired…

Computer Vision and Pattern Recognition · Computer Science 2025-07-25 Baoyao Yang , Wanyun Li , Dixin Chen , Junxiang Chen , Wenbin Yao , Haifeng Lin