English
Related papers

Related papers: MUGEN: A Playground for Video-Audio-Text Multimoda…

200 papers

We present \textsc{Vx2Text}, a framework for text generation from multimodal inputs consisting of video plus text, speech, or audio. In order to leverage transformer networks, which have been shown to be effective at modeling language, each…

Computer Vision and Pattern Recognition · Computer Science 2021-02-02 Xudong Lin , Gedas Bertasius , Jue Wang , Shih-Fu Chang , Devi Parikh , Lorenzo Torresani

Existing datasets for audio understanding primarily focus on single-turn interactions (i.e. audio captioning, audio question answering) for describing audio in natural language, thus limiting understanding audio via interactive dialogue. To…

Computation and Language · Computer Science 2024-04-12 Arushi Goel , Zhifeng Kong , Rafael Valle , Bryan Catanzaro

Video temporal understanding is crucial for multimodal large language models (MLLMs) to reason over events in videos. Despite recent advances in general video understanding, current MLLMs still struggle with fine-grained temporal reasoning.…

Computer Vision and Pattern Recognition · Computer Science 2026-04-21 Fuwen Luo , Shengfeng Lou , Chi Chen , Ziyue Wang , Chenliang Li , Weizhou Shen , Jiyue Guo , Peng Li , Ming Yan , Ji Zhang , Fei Huang , Yang Liu

Narrative videos, such as movies, pose significant challenges in video understanding due to their rich contexts (characters, dialogues, storylines) and diverse demands (identify who, relationship, and reason). In this paper, we introduce…

Computer Vision and Pattern Recognition · Computer Science 2024-09-13 Kevin Qinghong Lin , Pengchuan Zhang , Difei Gao , Xide Xia , Joya Chen , Ziteng Gao , Jinheng Xie , Xuhong Xiao , Mike Zheng Shou

We introduce QuerYD, a new large-scale dataset for retrieval and event localisation in video. A unique feature of our dataset is the availability of two audio tracks for each video: the original audio, and a high-quality spoken description…

Computer Vision and Pattern Recognition · Computer Science 2021-02-18 Andreea-Maria Oncescu , João F. Henriques , Yang Liu , Andrew Zisserman , Samuel Albanie

Characterizing and quantifying gender representation disparities in audiovisual storytelling contents is necessary to grasp how stereotypes may perpetuate on screen. In this article, we consider the high-level construct of objectification…

We propose a new task named Audio-driven Per-formance Video Generation (APVG), which aims to synthesizethe video of a person playing a certain instrument guided bya given music audio clip. It is a challenging task to gener-ate the…

Computer Vision and Pattern Recognition · Computer Science 2020-11-06 Hao Zhu , Yi Li , Feixia Zhu , Aihua Zheng , Ran He

Multimodal story customization aims to generate coherent story flows conditioned on textual descriptions, reference identity images, and shot types. While recent progress in story generation has shown promising results, most approaches rely…

Computer Vision and Pattern Recognition · Computer Science 2026-03-31 Wei-Hua Li , Cheng Sun , Chu-Song Chen

Multi-modal retrieval is an important problem for many applications, such as recommendation and search. Current benchmarks and even datasets are often manually constructed and consist of mostly clean samples where all modalities are…

Computer Vision and Pattern Recognition · Computer Science 2022-10-21 Laura Hanu , James Thewlis , Yuki M. Asano , Christian Rupprecht

Accurate emotion understanding in videos necessitates effectively recognizing and interpreting emotional states by integrating visual, textual, auditory, and contextual cues. Although recent Large Multimodal Models (LMMs) have exhibited…

Teasers are an effective tool for promoting content in entertainment, commercial and educational fields. However, creating an effective teaser for long videos is challenging for it requires long-range multimodal modeling on the input…

Computer Vision and Pattern Recognition · Computer Science 2024-11-12 Weihan Xu , Paul Pu Liang , Haven Kim , Julian McAuley , Taylor Berg-Kirkpatrick , Hao-Wen Dong

Multi-hop Question Generation (QG) effectively evaluates reasoning but remains confined to text; Video Question Generation (VideoQG) is limited to zero-hop questions over single segments. To address this, we introduce VideoChain, a novel…

Computer Vision and Pattern Recognition · Computer Science 2025-11-12 Arpan Phukan , Anupam Pandey , Deepjyoti Bodo , Asif Ekbal

In recent years, there have been significant advances in building end-to-end Machine Learning (ML) systems that learn at scale. But most of these systems are: (a) isolated (perception, speech, or language only); (b) trained on static…

Text-driven motion generation has attracted increasing attention due to its broad applications in virtual reality, animation, and robotics. While existing methods typically model human and animal motion separately, a joint cross-species…

Computer Vision and Pattern Recognition · Computer Science 2025-11-18 Xuan Wang , Kai Ruan , Liyang Qian , Zhizhi Guo , Chang Su , Gaoang Wang

Humans naturally understand moments in a video by integrating visual and auditory cues. For example, localizing a scene in the video like "A scientist passionately speaks on wildlife conservation as dramatic orchestral music plays, with the…

Computation and Language · Computer Science 2026-02-03 Zinuo Li , Xian Zhang , Yongxin Guo , Mohammed Bennamoun , Farid Boussaid , Girish Dwivedi , Luqi Gong , Qiuhong Ke

Music tagging and content-based retrieval systems have traditionally been constructed using pre-defined ontologies covering a rigid set of music attributes or text queries. This paper presents MuLan: a first attempt at a new generation of…

Audio and Speech Processing · Electrical Eng. & Systems 2022-08-29 Qingqing Huang , Aren Jansen , Joonseok Lee , Ravi Ganti , Judith Yue Li , Daniel P. W. Ellis

Text-to-image generation has achieved astonishing results, yet precise spatial controllability and prompt fidelity remain highly challenging. This limitation is typically addressed through cumbersome prompt engineering, scene layout…

Computer Vision and Pattern Recognition · Computer Science 2024-04-04 Petru-Daniel Tudosiu , Yongxin Yang , Shifeng Zhang , Fei Chen , Steven McDonagh , Gerasimos Lampouras , Ignacio Iacobacci , Sarah Parisot

Training multimodal large language models (MLLMs) for video understanding requires large-scale annotated data spanning diverse tasks such as object counting, question answering, and segmentation. However, collecting and annotating…

Computer Vision and Pattern Recognition · Computer Science 2026-04-15 Tanzila Rahman , Renjie Liao , Leonid Sigal

Automatically evaluating multimodal generation presents a significant challenge, as automated metrics often struggle to align reliably with human evaluation, especially for complex tasks that involve multiple modalities. To address this, we…

Artificial Intelligence · Computer Science 2025-05-26 Jihan Yao , Yushi Hu , Yujie Yi , Bin Han , Shangbin Feng , Guang Yang , Bingbing Wen , Ranjay Krishna , Lucy Lu Wang , Yulia Tsvetkov , Noah A. Smith , Banghua Zhu

Bridging vision and natural language is a longstanding goal in computer vision and multimedia research. While earlier works focus on generating a single-sentence description for visual content, recent works have studied paragraph…

Multimedia · Computer Science 2020-05-15 Junnan Li , Yongkang Wong , Qi Zhao , Mohan S. Kankanhalli