中文
相关论文

相关论文: MM-Narrator: Narrating Long-form Videos with Multi…

200 篇论文

Data visualizations and narratives are often integrated to convey data stories effectively. Among various data storytelling formats, data videos have been garnering increasing attention. These videos provide an intuitive interpretation of…

人机交互 · 计算机科学 2023-08-10 Leixian Shen , Yizhi Zhang , Haidong Zhang , Yun Wang

Multi-modal multi-party conversation (MMC) is a less studied yet important topic of research due to that it well fits real-world scenarios and thus potentially has more widely-used applications. Compared with the traditional multi-modal…

计算与语言 · 计算机科学 2024-12-24 Yueqian Wang , Xiaojun Meng , Yuxuan Wang , Jianxin Liang , Qun Liu , Dongyan Zhao

With the rapid development of foundation video generation technologies, long video generation models have exhibited promising research potential thanks to expanded content creation space. Recent studies reveal that the goal of long video…

计算机视觉与模式识别 · 计算机科学 2026-03-09 X. Feng , H. Yu , M. Wu , S. Hu , J. Chen , C. Zhu , J. Wu , X. Chu , K. Huang

Audio descriptions (ADs) function as acoustic commentaries designed to assist blind persons and persons with visual impairments in accessing digital media content on television and in movies, among other settings. As an accessibility…

计算与语言 · 计算机科学 2024-10-14 Yingqiang Gao , Lukas Fischer , Alexa Lintner , Sarah Ebling

Recent large multimodal models (LMMs) have become increasingly capable on image and video understanding, yet still struggle to sustain 4D continuous spatiotemporal dynamic reasoning. To study this capability gap, we formulate…

计算机视觉与模式识别 · 计算机科学 2026-05-20 Chaoyue Li , Yongxue Xu , Jie Feng , Jiayu Ding

There has been growing sentiment recently that modern large multimodal models (LMMs) have addressed most of the key challenges related to short video comprehension. As a result, both academia and industry are gradually shifting their…

计算机视觉与模式识别 · 计算机科学 2024-10-04 Jianrui Zhang , Mu Cai , Yong Jae Lee

The recent progress in Large Language Models (LLM) has spurred various advancements in image-language conversation agents, while how to build a proficient video-based dialogue system is still under exploration. Considering the extensive…

计算机视觉与模式识别 · 计算机科学 2024-06-28 Ruyang Liu , Chen Li , Yixiao Ge , Ying Shan , Thomas H. Li , Ge Li

Recent focus in video captioning has been on designing architectures that can consume both video and text modalities, and using large-scale video datasets with text transcripts for pre-training, such as HowTo100M. Though these approaches…

计算机视觉与模式识别 · 计算机科学 2023-06-23 Yuhan Shen , Linjie Yang , Longyin Wen , Haichao Yu , Ehsan Elhamifar , Heng Wang

To generate proper captions for videos, the inference needs to identify relevant concepts and pay attention to the spatial relationships between them as well as to the temporal development in the clip. Our end-to-end encoder-decoder video…

计算机视觉与模式识别 · 计算机科学 2022-08-22 Zohreh Ghaderi , Leonard Salewski , Hendrik P. A. Lensch

Long-form video understanding remains a fundamental challenge for current Video Large Language Models. Most existing models rely on static reasoning over uniformly sampled frames, which weakens temporal localization and leads to substantial…

计算机视觉与模式识别 · 计算机科学 2026-04-21 Chenglin Li , Qianglong Chen , Feng Han , Yikun Wang , Xingxi Yin , Yan Gong , Ruilin Li , Yin Zhang , Jiaqi Wang

Generating Audio Description (AD) for movies is a challenging task that requires fine-grained visual understanding and an awareness of the characters and their names. Currently, visual language models for AD generation are limited by a lack…

计算机视觉与模式识别 · 计算机科学 2024-04-23 Tengda Han , Max Bain , Arsha Nagrani , Gül Varol , Weidi Xie , Andrew Zisserman

Multimodal Large Language Models (MLLMs) perform well in video understanding but degrade on long videos due to fixed-length context and weak long-term dependency modeling. Retrieval-Augmented Generation (RAG) can expand knowledge…

计算机视觉与模式识别 · 计算机科学 2025-11-25 Zhucun Xue , Jiangning Zhang , Xurong Xie , Yuxuan Cai , Yong Liu , Xiangtai Li , Dacheng Tao

Advertisement videos (ads) play an integral part in the domain of Internet e-commerce as they amplify the reach of particular products to a broad audience or can serve as a medium to raise awareness about specific issues through concise…

计算机视觉与模式识别 · 计算机科学 2023-08-29 Digbalay Bose , Rajat Hebbar , Tiantian Feng , Krishna Somandepalli , Anfeng Xu , Shrikanth Narayanan

Despite significant advancements in Large Language Models (LLMs) and Large Vision-Language Models (LVLMs), current models still face substantial challenges in handling complex, multi-turn, and visually-grounded tasks that demand deep…

计算与语言 · 计算机科学 2025-08-22 Seungmin Han , Haeun Kwon , Ji-jun Park , Taeyang Yoon

This paper focuses on the challenge of answering questions in scenarios that are composed of rich and complex dynamic audio-visual components. Although existing Multimodal Large Language Models (MLLMs) can respond to audio-visual content,…

计算机视觉与模式识别 · 计算机科学 2024-03-08 Qilang Ye , Zitong Yu , Rui Shao , Xinyu Xie , Philip Torr , Xiaochun Cao

Large Language Models (LLMs) have achieved remarkable reliability and advanced capabilities through extended test-time reasoning. However, extending these capabilities to Multi-modal Large Language Models (MLLMs) remains a significant…

计算机视觉与模式识别 · 计算机科学 2026-03-20 Yuhao Dong , Zuyan Liu , Shulin Tian , Yongming Rao , Ziwei Liu

Large Multimodal Models (LMMs) have demonstrated exceptional performance across a wide range of domains. This paper explores their potential in pronunciation assessment tasks, with a particular focus on evaluating the capabilities of the…

声音 · 计算机科学 2025-03-17 Ke Wang , Lei He , Kun Liu , Yan Deng , Wenning Wei , Sheng Zhao

Employing Multimodal Large Language Models (MLLMs) for long video understanding remains a challenging problem due to the dilemma between the substantial number of video frames (i.e., visual tokens) versus the limited context length of…

计算机视觉与模式识别 · 计算机科学 2025-06-02 Bo Fang , Wenhao Wu , Qiangqiang Wu , Yuxin Song , Antoni B. Chan

We propose the first joint audio-video generation framework that brings engaging watching and listening experiences simultaneously, towards high-quality realistic videos. To generate joint audio-video pairs, we propose a novel Multi-Modal…

计算机视觉与模式识别 · 计算机科学 2023-03-27 Ludan Ruan , Yiyang Ma , Huan Yang , Huiguo He , Bei Liu , Jianlong Fu , Nicholas Jing Yuan , Qin Jin , Baining Guo

In this paper, we propose and design a new task called audio moment retrieval (AMR). Unlike conventional language-based audio retrieval tasks that search for short audio clips from an audio database, AMR aims to predict relevant moments in…

音频与语音处理 · 电气工程与系统科学 2025-08-05 Hokuto Munakata , Taichi Nishimura , Shota Nakada , Tatsuya Komatsu