中文
相关论文

相关论文: Paper2Video: Automatic Video Generation from Scien…

200 篇论文

The dissemination of scholarly research is critical, yet researchers often lack the time and skills to create engaging content for popular media such as short-form videos. To address this gap, we explore the use of generative AI to help…

The paper-to-video task converts a research paper into a structured video abstract, distilling key concepts, methods, and conclusions into an accessible, well-organized format. While state-of-the-art video generation models demonstrate…

计算机视觉与模式识别 · 计算机科学 2025-09-09 Jingwei Liu , Ling Yang , Hao Luo , Fan Wang , Hongyan Li , Mengdi Wang

Creating presentation materials requires complex multimodal reasoning skills to summarize key concepts and arrange them in a logical and visually pleasing manner. Can machines learn to emulate this laborious process? We present a novel task…

计算机视觉与模式识别 · 计算机科学 2022-03-22 Tsu-Jui Fu , William Yang Wang , Daniel McDuff , Yale Song

Currently, no large-scale training data is available for the task of scientific paper summarization. In this paper, we propose a novel method that automatically generates summaries for scientific papers, by utilizing videos of talks at…

计算与语言 · 计算机科学 2019-06-14 Guy Lev , Michal Shmueli-Scheuer , Jonathan Herzig , Achiya Jerbi , David Konopnicki

Academic project websites can more effectively disseminate research when they clearly present core content and enable intuitive navigation and interaction. However, current approaches such as direct Large Language Model (LLM) generation,…

计算与语言 · 计算机科学 2025-10-20 Yuhang Chen , Tianpeng Lv , Siyi Zhang , Yixiang Yin , Yao Wan , Philip S. Yu , Dongping Chen

Academic poster generation is a crucial yet challenging task in scientific communication, requiring the compression of long-context interleaved documents into a single, visually coherent page. To address this challenge, we introduce the…

计算机视觉与模式识别 · 计算机科学 2025-10-31 Wei Pang , Kevin Qinghong Lin , Xiangru Jian , Xi He , Philip Torr

Automatic presentation slide generation can greatly streamline content creation. However, since preferences of each user may vary, existing under-specified formulations often lead to suboptimal results that fail to align with individual…

计算与语言 · 计算机科学 2025-12-24 Wenzheng Zeng , Mingyu Ouyang , Langyuan Cui , Hwee Tou Ng

Despite the rapid growth of machine learning research, corresponding code implementations are often unavailable, making it slow and labor-intensive for researchers to reproduce results and build upon prior work. In the meantime, recent…

计算与语言 · 计算机科学 2026-03-02 Minju Seo , Jinheon Baek , Seongyun Lee , Sung Ju Hwang

While recent generative models advance pixel-space video synthesis, they remain limited in producing professional educational videos, which demand disciplinary knowledge, precise visual structures, and coherent transitions, limiting their…

计算机视觉与模式识别 · 计算机科学 2025-10-02 Yanzhe Chen , Kevin Qinghong Lin , Mike Zheng Shou

Automating the creation of scientific diagrams from academic papers can significantly streamline the development of tutorials, presentations, and posters, thereby saving time and accelerating the process. Current text-to-image models…

计算与语言 · 计算机科学 2024-10-17 Ishani Mondal , Zongxia Li , Yufang Hou , Anandhavelu Natarajan , Aparna Garimella , Jordan Boyd-Graber

Research papers are well structured documents. They have text, figures, equations, tables etc., to covey their ideas and findings. They are divided into sections like Introduction, Model, Experiments etc., which deal with different aspects…

计算与语言 · 计算机科学 2024-11-28 Keshav Kumar , Ravindranath Chowdary

We present PresentAgent, a multimodal agent that transforms long-form documents into narrated presentation videos. While existing approaches are limited to generating static slides or text summaries, our method advances beyond these…

计算机视觉与模式识别 · 计算机科学 2025-07-08 Jingwei Shi , Zeyu Zhang , Biao Wu , Yanjie Liang , Meng Fang , Ling Chen , Yang Zhao

Generating academic slides from scientific papers is a challenging multimodal reasoning task that requires both long context understanding and deliberate visual planning. Existing approaches largely reduce it to text only summarization,…

人工智能 · 计算机科学 2025-12-10 Xin Liang , Xiang Zhang , Yiwei Xu , Siqi Sun , Chenyu You

Research consumption has been traditionally limited to the reading of academic papers-a static, dense, and formally written format. Alternatively, pre-recorded conference presentation videos, which are more dynamic, concise, and colloquial,…

人机交互 · 计算机科学 2023-08-30 Tae Soo Kim , Matt Latzke , Jonathan Bragg , Amy X. Zhang , Joseph Chee Chang

Transforming scientific papers into multimodal presentation content is essential for research dissemination but remains labor intensive. Existing automated solutions typically treat each format as an isolated downstream task, leading to…

Academic posters are vital for scholarly communication, yet their manual creation is time-consuming. However, automated academic poster generation faces significant challenges in preserving intricate scientific details and achieving…

计算与语言 · 计算机科学 2025-05-26 Tao Sun , Enhao Pan , Zhengkai Yang , Kaixin Sui , Jiajun Shi , Xianfu Cheng , Tongliang Li , Wenhao Huang , Ge Zhang , Jian Yang , Zhoujun Li

The technical complexity of research papers often limits their reach, necessitating more accessible formats like scientific videos to disseminate key insights through engaging narration. However, existing automated methods primarily focus…

人工智能 · 计算机科学 2026-04-22 Xiao Liang , Bangxin Li , Zixuan Chen , Hanyue Zheng , Zhi Ma , Di Wang , Cong Tian , Quan Wang

Automatically generating presentations from documents is a challenging task that requires accommodating content quality, visual appeal, and structural coherence. Existing methods primarily focus on improving and evaluating the content…

人工智能 · 计算机科学 2025-02-24 Hao Zheng , Xinyan Guan , Hao Kong , Jia Zheng , Weixiang Zhou , Hongyu Lin , Yaojie Lu , Ben He , Xianpei Han , Le Sun

Generating presentation slides is a time-consuming task that urgently requires automation. Due to their limited flexibility and lack of automated refinement mechanisms, existing autonomous LLM-based agents face constraints in real-world…

计算与语言 · 计算机科学 2025-02-24 Yunqing Xu , Xinbei Ma , Jiyang Qiu , Hai Zhao

Video generation has achieved remarkable progress, with generated videos increasingly resembling real ones. However, the rapid advance in generation has outpaced the development of adequate evaluation metrics. Currently, the assessment of…

计算机视觉与模式识别 · 计算机科学 2026-05-21 Nabyl Quignon , Baptiste Chopin , Yaohui Wang , Antitza Dantcheva
‹ 上一页 1 2 3 10 下一页 ›