中文
相关论文

相关论文: VideoAgent: Personalized Synthesis of Scientific V…

200 篇论文

We present PresentAgent, a multimodal agent that transforms long-form documents into narrated presentation videos. While existing approaches are limited to generating static slides or text summaries, our method advances beyond these…

计算机视觉与模式识别 · 计算机科学 2025-07-08 Jingwei Shi , Zeyu Zhang , Biao Wu , Yanjie Liang , Meng Fang , Ling Chen , Yang Zhao

Video Question Answering (VQA) inherently relies on multimodal reasoning, integrating visual, temporal, and linguistic cues to achieve a deeper understanding of video content. However, many existing methods rely on feeding frame-level…

Recent advancements in multimodal large language models (MLLMs) and video agent systems have significantly improved general video understanding. However, when applied to scientific video understanding and educating, a domain that demands…

计算机视觉与模式识别 · 计算机科学 2025-11-25 Zhiyu Xu , Weilong Yan , Yufei Shi , Xin Meng , Tao He , Huiping Zhuang , Ming Li , Hehe Fan

Video question answering (VideoQA) is a challenging task that requires integrating spatial, temporal, and semantic information to capture the complex dynamics of video sequences. Although recent advances have introduced various approaches…

计算机视觉与模式识别 · 计算机科学 2026-04-08 Zhongyu Yang , Zuhao Yang , Shuo Zhan , Tan Yue , Wei Pang , Yingfang Yuan

With the advancement of AIGC (AI-generated content) technologies, an increasing number of generative models are revolutionizing fields such as video editing, music generation, and even film production. However, due to the limitations of…

计算机视觉与模式识别 · 计算机科学 2026-01-07 Daoan Zhang , Wenlin Yao , Xiaoyang Wang , Yebowen Hu , Jiebo Luo , Dong Yu

Presentation generation is moving beyond static slide creation toward end-to-end presentation video generation with research grounding, multimodal media, and interactive delivery. We introduce PresentAgent-2, an agentic framework for…

计算机视觉与模式识别 · 计算机科学 2026-05-13 Wei Wu , Ziyang Xu , Zeyu Zhang , Yang Zhao , Hao Tang

Generating academic slides from scientific papers is a challenging multimodal reasoning task that requires both long context understanding and deliberate visual planning. Existing approaches largely reduce it to text only summarization,…

人工智能 · 计算机科学 2025-12-10 Xin Liang , Xiang Zhang , Yiwei Xu , Siqi Sun , Chenyu You

We explore how reconciling several foundation models (large language models and vision-language models) with a novel unified memory mechanism could tackle the challenging video understanding problem, especially capturing the long-term…

计算机视觉与模式识别 · 计算机科学 2024-07-16 Yue Fan , Xiaojian Ma , Rujie Wu , Yuntao Du , Jiaqi Li , Zhi Gao , Qing Li

Video generation has been used to generate visual plans for controlling robotic systems. Given an image observation and a language instruction, previous work has generated video plans which are then converted to robot controls to be…

Facing scaling laws, video data from the internet becomes increasingly important. However, collecting extensive videos that meet specific needs is extremely labor-intensive and time-consuming. In this work, we study the way to expedite this…

人工智能 · 计算机科学 2025-09-26 Yidan Zhang , Mutian Xu , Yiming Hao , Kun Zhou , Jiahao Chang , Xiaoqiang Liu , Pengfei Wan , Hongbo Fu , Xiaoguang Han

Story visualization is the transformation of narrative elements into image sequences. While existing research has primarily focused on visual contextual coherence, the deeper narrative essence of stories often remains overlooked. This…

计算机视觉与模式识别 · 计算机科学 2025-03-05 Seungkwon Kim , GyuTae Park , Sangyeon Kim , Seung-Hun Nam

Presentation slides are a primary medium for data-driven reporting, yet keeping complex, analytics-style decks up to date remains labor-intensive. Existing automation methods mostly follow fixed template filling and cannot support dynamic…

计算与语言 · 计算机科学 2026-04-21 Kun Zhou , Jiakai He , Wenmian Yang , Zhensheng Wang , Yiquan Zhang , Weijia Jia

Video-to-audio synthesis, which generates synchronized audio for visual content, critically enhances viewer immersion and narrative coherence in film and interactive media. However, video-to-audio dubbing for long-form content remains an…

计算机视觉与模式识别 · 计算机科学 2025-03-18 Yehang Zhang , Xinli Xu , Xiaojie Xu , Li Liu , Yingcong Chen

Academic presentation videos have become an essential medium for research communication, yet producing them remains highly labor-intensive, often requiring hours of slide design, recording, and editing for a short 2 to 10 minutes video.…

计算机视觉与模式识别 · 计算机科学 2025-10-10 Zeyu Zhu , Kevin Qinghong Lin , Mike Zheng Shou

Web agents struggle to adapt to new websites due to the scarcity of environment specific tasks and demonstrations. Recent works have explored synthetic data generation to address this challenge, however, they suffer from data quality issues…

The advent of AI-Generated Content (AIGC) has spurred research into automated video generation to streamline conventional processes. However, automating storytelling video production, particularly for customized narratives, remains…

计算机视觉与模式识别 · 计算机科学 2024-11-12 Panwen Hu , Jin Jiang , Jianqi Chen , Mingfei Han , Shengcai Liao , Xiaojun Chang , Xiaodan Liang

Generating engaging, accurate short-form videos from scientific papers is challenging due to content complexity and the gap between expert authors and readers. Existing end-to-end methods often suffer from factual inaccuracies and visual…

计算与语言 · 计算机科学 2025-04-29 Jong Inn Park , Maanas Taneja , Qianwen Wang , Dongyeop Kang

We introduce V-Agent, a novel multi-agent platform designed for advanced video search and interactive user-system conversations. By fine-tuning a vision-language model (VLM) with a small video preference dataset and enhancing it with a…

计算机视觉与模式识别 · 计算机科学 2026-01-08 SunYoung Park , Jong-Hyeon Lee , Youngjune Kim , Daegyu Sung , Younghyun Yu , Young-rok Cha , Jeongho Ju

Omnimodal large language models have made significant strides in unifying audio and visual modalities; however, they often face challenges in fine-grained cross-modal understanding and have difficulty with multimodal alignment. To address…

计算机视觉与模式识别 · 计算机科学 2026-02-06 Keda Tao , Wenjie Du , Bohan Yu , Weiqiang Wang , Jian Liu , Huan Wang

The rapid advancement of video generation has rendered existing evaluation systems inadequate for assessing state-of-the-art models, primarily due to simple prompts that cannot showcase the model's capabilities, fixed evaluation operators…

计算机视觉与模式识别 · 计算机科学 2025-04-29 Yuhang Yang , Ke Fan , Shangkun Sun , Hongxiang Li , Ailing Zeng , FeiLin Han , Wei Zhai , Wei Liu , Yang Cao , Zheng-Jun Zha
‹ 上一页 1 2 3 10 下一页 ›