中文
相关论文

相关论文: Multi-sentence Video Grounding for Long Video Gene…

200 篇论文

Given an untrimmed video and natural language query, video sentence grounding aims to localize the target temporal moment in the video. Existing methods mainly tackle this task by matching and aligning semantics of the descriptive sentence…

计算机视觉与模式识别 · 计算机科学 2022-12-29 Wei Ji , Long Chen , Yinwei Wei , Yiming Wu , Tat-Seng Chua

Video storytelling is engaging multimedia content that utilizes video and its accompanying narration to attract the audience, where a key challenge is creating narrations for recorded visual scenes. Previous studies on dense video…

多媒体 · 计算机科学 2024-12-31 Dingyi Yang , Chunru Zhan , Ziheng Wang , Biao Wang , Tiezheng Ge , Bo Zheng , Qin Jin

Human motion generation is a significant pursuit in generative computer vision with widespread applications in film-making, video games, AR/VR, and human-robot interaction. Current methods mainly utilize either diffusion-based generative…

计算机视觉与模式识别 · 计算机科学 2025-02-03 Canxuan Gang

Storytelling video generation (SVG) aims to produce coherent and visually rich multi-scene videos that follow a structured narrative. Existing methods primarily employ LLM for high-level planning to decompose a story into scene-level…

计算机视觉与模式识别 · 计算机科学 2025-11-17 Zun Wang , Jialu Li , Han Lin , Jaehong Yoon , Mohit Bansal

We present Movie Gen, a cast of foundation models that generates high-quality, 1080p HD videos with different aspect ratios and synchronized audio. We also show additional capabilities such as precise instruction-based video editing and…

Recent advances in large-scale pre-training such as GPT-3 allow seemingly high quality text to be generated from a given prompt. However, such generation systems often suffer from problems of hallucinated facts, and are not inherently…

计算与语言 · 计算机科学 2022-02-25 Yizhe Zhang , Siqi Sun , Xiang Gao , Yuwei Fang , Chris Brockett , Michel Galley , Jianfeng Gao , Bill Dolan

Advances in technology have led to the development of methods that can create desired visual multimedia. In particular, image generation using deep learning has been extensively studied across diverse fields. In comparison, video…

计算机视觉与模式识别 · 计算机科学 2021-06-29 Doyeon Kim , Donggyu Joo , Junmo Kim

Spatio-temporal consistency is a critical research topic in video generation. A qualified generated video segment must ensure plot plausibility and coherence while maintaining visual consistency of objects and scenes across varying…

计算机视觉与模式识别 · 计算机科学 2025-03-11 Runze Zhang , Guoguang Du , Xiaochuan Li , Qi Jia , Liang Jin , Lu Liu , Jingjing Wang , Cong Xu , Zhenhua Guo , Yaqian Zhao , Xiaoli Gong , Rengang Li , Baoyu Fan

Short videos are an effective tool for promoting contents and improving knowledge accessibility. While existing extractive video summarization methods struggle to produce a coherent narrative, existing abstractive methods cannot `quote'…

计算机视觉与模式识别 · 计算机科学 2025-05-27 Weihan Xu , Yimeng Ma , Jingyue Huang , Yang Li , Wenye Ma , Taylor Berg-Kirkpatrick , Julian McAuley , Paul Pu Liang , Hao-Wen Dong

Long-context video modeling is essential for enabling generative models to function as world simulators, as they must maintain temporal coherence over extended time spans. However, most existing models are trained on short clips, limiting…

计算机视觉与模式识别 · 计算机科学 2025-05-20 Yuchao Gu , Weijia Mao , Mike Zheng Shou

Existing text-to-video (T2V) models often struggle with generating videos with sufficiently pronounced or complex actions. A key limitation lies in the text prompt's inability to precisely convey intricate motion details. To address this,…

计算机视觉与模式识别 · 计算机科学 2024-11-14 Qiang Zhou , Shaofeng Zhang , Nianzu Yang , Ye Qian , Hao Li

Long-term Video Question Answering (VideoQA) is a challenging vision-and-language bridging task focusing on semantic understanding of untrimmed long-term videos and diverse free-form questions, simultaneously emphasizing comprehensive…

计算机视觉与模式识别 · 计算机科学 2024-10-15 Ting Yu , Kunhao Fu , Jian Zhang , Qingming Huang , Jun Yu

Temporal sentence grounding in videos aims to detect and localize one target video segment, which semantically corresponds to a given sentence. Existing methods mainly tackle this task via matching and aligning semantics between a sentence…

计算机视觉与模式识别 · 计算机科学 2019-11-01 Yitian Yuan , Lin Ma , Jingwen Wang , Wei Liu , Wenwu Zhu

In this work, we tackle the problem of long-form video-language grounding (VLG). Given a long-form video and a natural language query, a model should temporally localize the precise moment that answers the query. Humans can easily solve VLG…

计算机视觉与模式识别 · 计算机科学 2024-08-07 Hyogun Lee , Soyeon Hong , Mujeen Sung , Jinwoo Choi

This paper considers the problem of Multi-Hop Video Question Answering (MH-VidQA) in long-form egocentric videos. This task not only requires to answer visual questions, but also to localize multiple relevant time intervals within the video…

计算机视觉与模式识别 · 计算机科学 2024-08-27 Qirui Chen , Shangzhe Di , Weidi Xie

In long-video understanding, conventional uniform frame sampling often fails to capture key visual evidence, leading to degraded performance and increased hallucinations. To address this, recent agentic thinking-with-videos paradigms have…

Large-scale language models (LMs) pretrained on massive corpora of text, such as GPT-2, are powerful open-domain text generators. However, as our systematic examination reveals, it is still challenging for such models to generate coherent…

计算与语言 · 计算机科学 2021-04-15 Bowen Tan , Zichao Yang , Maruan AI-Shedivat , Eric P. Xing , Zhiting Hu

Recent video diffusion models have enhanced video editing, but it remains challenging to handle instructional editing and diverse tasks (e.g., adding, removing, changing) within a unified framework. In this paper, we introduce VEGGIE, a…

计算机视觉与模式识别 · 计算机科学 2025-10-28 Shoubin Yu , Difan Liu , Ziqiao Ma , Yicong Hong , Yang Zhou , Hao Tan , Joyce Chai , Mohit Bansal

Generating temporally coherent, long-duration videos with precise control over subject identity and movement remains a fundamental challenge for contemporary diffusion-based models, which often suffer from identity drift and are limited to…

计算机视觉与模式识别 · 计算机科学 2026-04-13 Jingxuan He , Busheng Su , Finn Wong

Multimodal large language models (MLLMs) have achieved remarkable progress in video understanding. However, seemingly plausible outputs often suffer from poor visual and temporal grounding: a model may fabricate object existence, assign…

计算机视觉与模式识别 · 计算机科学 2026-04-03 Yihao Quan , Zeru Shi , Jinman Zhao , Ruixiang Tang