中文
相关论文

相关论文: A Versatile Multimodal Agent for Multimedia Conten…

200 篇论文

We present PresentAgent, a multimodal agent that transforms long-form documents into narrated presentation videos. While existing approaches are limited to generating static slides or text summaries, our method advances beyond these…

计算机视觉与模式识别 · 计算机科学 2025-07-08 Jingwei Shi , Zeyu Zhang , Biao Wu , Yanjie Liang , Meng Fang , Ling Chen , Yang Zhao

The rapid progress of Artificial Intelligence Generated Content (AIGC) tools enables images, videos, and visualizations to be created on demand for webpage design, offering a flexible and increasingly adopted paradigm for modern UI/UX.…

计算机视觉与模式识别 · 计算机科学 2026-04-17 Yan Li , Zezi Zeng , Yifan Yang , Yuqing Yang , Ning Liao , Weiwei Guo , Lili Qiu , Mingxi Cheng , Qi Dai , Zhendong Wang , Zhengyuan Yang , Xue Yang , Ji Li , Lijuan Wang , Chong Luo

AI-generated content (AIGC) methods aim to produce text, images, videos, 3D assets, and other media using AI algorithms. Due to its wide range of applications and the potential of recent works, AIGC developments -- especially in Machine…

计算机视觉与模式识别 · 计算机科学 2025-01-22 Lin Geng Foo , Hossein Rahmani , Jun Liu

Presentation generation is moving beyond static slide creation toward end-to-end presentation video generation with research grounding, multimodal media, and interactive delivery. We introduce PresentAgent-2, an agentic framework for…

计算机视觉与模式识别 · 计算机科学 2026-05-13 Wei Wu , Ziyang Xu , Zeyu Zhang , Yang Zhao , Hao Tang

While AI excels at generating text, audio, images, and videos, creating interactive audio-visual content such as video games remains challenging. Current LLMs can generate JavaScript games and animations, but lack automated evaluation…

人工智能 · 计算机科学 2025-08-04 Alexia Jolicoeur-Martineau

The rapid advancement of video generation has rendered existing evaluation systems inadequate for assessing state-of-the-art models, primarily due to simple prompts that cannot showcase the model's capabilities, fixed evaluation operators…

计算机视觉与模式识别 · 计算机科学 2025-04-29 Yuhang Yang , Ke Fan , Shangkun Sun , Hongxiang Li , Ailing Zeng , FeiLin Han , Wei Zhai , Wei Liu , Yang Cao , Zheng-Jun Zha

The Agent and AIGC (Artificial Intelligence Generated Content) technologies have recently made significant progress. We propose AesopAgent, an Agent-driven Evolutionary System on Story-to-Video Production. AesopAgent is a practical…

计算机视觉与模式识别 · 计算机科学 2024-03-14 Jiuniu Wang , Zehua Du , Yuyuan Zhao , Bo Yuan , Kexiang Wang , Jian Liang , Yaxi Zhao , Yihen Lu , Gengliang Li , Junlong Gao , Xin Tu , Zhenyu Guo

With the success of 2D diffusion models, 2D AIGC content has already transformed our lives. Recently, this success has been extended to 3D AIGC, with state-of-the-art methods generating textured 3D models from single images or text.…

计算机视觉与模式识别 · 计算机科学 2024-12-19 Junhao Chen , Xiang Li , Xiaojun Ye , Chao Li , Zhaoxin Fan , Hao Zhao

Large language models (LLMs) have enabled remarkable advances in automated task-solving with multi-agent systems. However, most existing LLM-based multi-agent approaches rely on predefined agents to handle simple tasks, limiting the…

人工智能 · 计算机科学 2024-05-01 Guangyao Chen , Siwei Dong , Yu Shu , Ge Zhang , Jaward Sesay , Börje F. Karlsson , Jie Fu , Yemin Shi

Multimodality-to-Multiaudio (MM2MA) generation faces significant challenges in synthesizing diverse and contextually aligned audio types (e.g., sound effects, speech, music, and songs) from multimodal inputs (e.g., video, text, images),…

声音 · 计算机科学 2025-08-06 Yan Rong , Jinting Wang , Guangzhi Lei , Shan Yang , Li Liu

This paper proposes a multi-agent artificial intelligence system that generates response-oriented media content in real time based on audio-derived emotional signals. Unlike conventional speech emotion recognition studies that focus…

人工智能 · 计算机科学 2026-01-21 HyeYoung Lee

AI-empowered music processing is a diverse field that encompasses dozens of tasks, ranging from generation tasks (e.g., timbre synthesis) to comprehension tasks (e.g., music classification). For developers and amateurs, it is very difficult…

计算与语言 · 计算机科学 2023-10-26 Dingyao Yu , Kaitao Song , Peiling Lu , Tianyu He , Xu Tan , Wei Ye , Shikun Zhang , Jiang Bian

The rapid advancement of large language models (LLMs) and artificial intelligence-generated content (AIGC) has accelerated AI-native applications, such as AI-based storybooks that automate engaging story production for children. However,…

计算与语言 · 计算机科学 2025-03-10 Xuenan Xu , Jiahao Mei , Chenliang Li , Yuning Wu , Ming Yan , Shaopeng Lai , Ji Zhang , Mengyue Wu

Recently, ChatGPT, along with DALL-E-2 and Codex,has been gaining significant attention from society. As a result, many individuals have become interested in related resources and are seeking to uncover the background and secrets behind its…

人工智能 · 计算机科学 2023-03-09 Yihan Cao , Siyu Li , Yixin Liu , Zhiling Yan , Yutong Dai , Philip S. Yu , Lichao Sun

The advent of AI-Generated Content (AIGC) has spurred research into automated video generation to streamline conventional processes. However, automating storytelling video production, particularly for customized narratives, remains…

计算机视觉与模式识别 · 计算机科学 2024-11-12 Panwen Hu , Jin Jiang , Jianqi Chen , Mingfei Han , Shengcai Liao , Xiaojun Chang , Xiaodan Liang

Despite rapid advancements in video generation models, generating coherent storytelling videos that span multiple scenes and characters remains challenging. Current methods often rigidly convert pre-generated keyframes into fixed-length…

多智能体系统 · 计算机科学 2025-10-03 Haoyuan Shi , Yunxin Li , Xinyu Chen , Longyue Wang , Baotian Hu , Min Zhang

Multi-modal AI systems will likely become a ubiquitous presence in our everyday lives. A promising approach to making these systems more interactive is to embody them as agents within physical and virtual environments. At present, systems…

The proliferation of generative AI systems creates unprecedented opportunities for content creation while raising critical concerns about controllability, copyright infringement, and content provenance. Current generative models operate as…

密码学与安全 · 计算机科学 2026-01-13 Haris Khan , Sadia Asif , Shumaila Asif

The technical complexity of research papers often limits their reach, necessitating more accessible formats like scientific videos to disseminate key insights through engaging narration. However, existing automated methods primarily focus…

人工智能 · 计算机科学 2026-04-22 Xiao Liang , Bangxin Li , Zixuan Chen , Hanyue Zheng , Zhi Ma , Di Wang , Cong Tian , Quan Wang

Video Question Answering (VQA) inherently relies on multimodal reasoning, integrating visual, temporal, and linguistic cues to achieve a deeper understanding of video content. However, many existing methods rely on feeding frame-level…

‹ 上一页 1 2 3 10 下一页 ›