中文
相关论文

相关论文: Multi-Agent Game Generation and Evaluation via Aud…

200 篇论文

The rapid advancement of video generation has rendered existing evaluation systems inadequate for assessing state-of-the-art models, primarily due to simple prompts that cannot showcase the model's capabilities, fixed evaluation operators…

计算机视觉与模式识别 · 计算机科学 2025-04-29 Yuhang Yang , Ke Fan , Shangkun Sun , Hongxiang Li , Ailing Zeng , FeiLin Han , Wei Zhai , Wei Liu , Yang Cao , Zheng-Jun Zha

With the advancement of AIGC (AI-generated content) technologies, an increasing number of generative models are revolutionizing fields such as video editing, music generation, and even film production. However, due to the limitations of…

计算机视觉与模式识别 · 计算机科学 2026-01-07 Daoan Zhang , Wenlin Yao , Xiaoyang Wang , Yebowen Hu , Jiebo Luo , Dong Yu

Music-to-Video (M2V) generation for full-length songs faces significant challenges. Existing methods produce short, disjointed clips, failing to align visuals with musical structure, beats, or lyrics, and lack temporal consistency. We…

Despite rapid advancements in video generation models, generating coherent storytelling videos that span multiple scenes and characters remains challenging. Current methods often rigidly convert pre-generated keyframes into fixed-length…

多智能体系统 · 计算机科学 2025-10-03 Haoyuan Shi , Yunxin Li , Xinyu Chen , Longyue Wang , Baotian Hu , Min Zhang

Cross-Video Reasoning (CVR) has emerged as a critical frontier in multimodal intelligence, requiring models to retrieve, align, and aggregate evidence distributed across multiple videos. Current Multimodal Large Language Models (MLLMs)…

计算机视觉与模式识别 · 计算机科学 2026-05-29 Yilun Qiu , Jiahe Wang , Cilin Yan , Jiayin Cai , Xiaolong Jiang , Yao Hu , Chun Yuan

Modern businesses are increasingly challenged by the time and expense required to generate and assess high-quality content. Human writers face time constraints, and extrinsic evaluations can be costly. While Large Language Models (LLMs)…

人工智能 · 计算机科学 2025-12-10 Thanh Vu , Richi Nayak , Thiru Balasubramaniam

Text-to-video generation has made significant strides, but replicating the capabilities of advanced systems like OpenAI Sora remains challenging due to their closed-source nature. Existing open-source methods struggle to achieve comparable…

计算机视觉与模式识别 · 计算机科学 2024-10-07 Zhengqing Yuan , Yixin Liu , Yihan Cao , Weixiang Sun , Haolong Jia , Ruoxi Chen , Zhaoxu Li , Bin Lin , Li Yuan , Lifang He , Chi Wang , Yanfang Ye , Lichao Sun

The rapid advancement of large language models (LLMs) and artificial intelligence-generated content (AIGC) has accelerated AI-native applications, such as AI-based storybooks that automate engaging story production for children. However,…

计算与语言 · 计算机科学 2025-03-10 Xuenan Xu , Jiahao Mei , Chenliang Li , Yuning Wu , Ming Yan , Shaopeng Lai , Ji Zhang , Mengyue Wu

We introduce Audio-Agent, a multimodal framework for audio generation, editing and composition based on text or video inputs. Conventional approaches for text-to-audio (TTA) tasks often make single-pass inferences from text descriptions.…

声音 · 计算机科学 2025-01-15 Zixuan Wang , Chi-Keung Tang , Yu-Wing Tai

We introduce V-Agent, a novel multi-agent platform designed for advanced video search and interactive user-system conversations. By fine-tuning a vision-language model (VLM) with a small video preference dataset and enhancing it with a…

计算机视觉与模式识别 · 计算机科学 2026-01-08 SunYoung Park , Jong-Hyeon Lee , Youngjune Kim , Daegyu Sung , Younghyun Yu , Young-rok Cha , Jeongho Ju

We present PresentAgent, a multimodal agent that transforms long-form documents into narrated presentation videos. While existing approaches are limited to generating static slides or text summaries, our method advances beyond these…

计算机视觉与模式识别 · 计算机科学 2025-07-08 Jingwei Shi , Zeyu Zhang , Biao Wu , Yanjie Liang , Meng Fang , Ling Chen , Yang Zhao

This paper proposes a multi-agent artificial intelligence system that generates response-oriented media content in real time based on audio-derived emotional signals. Unlike conventional speech emotion recognition studies that focus…

人工智能 · 计算机科学 2026-01-21 HyeYoung Lee

Although recent end-to-end video generation models demonstrate impressive performance in visually oriented content creation, they remain limited in scenarios that require strict logical rigor and precise knowledge representation, such as…

人工智能 · 计算机科学 2026-02-13 Lingyong Yan , Jiulong Wu , Dong Xie , Weixian Shi , Deguo Xia , Jizhou Huang

The advent of AI-Generated Content (AIGC) has spurred research into automated video generation to streamline conventional processes. However, automating storytelling video production, particularly for customized narratives, remains…

计算机视觉与模式识别 · 计算机科学 2024-11-12 Panwen Hu , Jin Jiang , Jianqi Chen , Mingfei Han , Shengcai Liao , Xiaojun Chang , Xiaodan Liang

Real-world visualization tasks involve complex, multi-modal requirements that extend beyond simple text-to-chart generation, requiring reference images, code examples, and iterative refinement. Current systems exhibit fundamental…

计算与语言 · 计算机科学 2026-01-27 Jinwei Lu , Yuanfeng Song , Chen Zhang , Raymond Chi-Wing Wong

Vision-language models (VLMs) have shown impressive capabilities in perceptual tasks, yet they degrade in complex multi-hop reasoning under multiplayer game settings with imperfect and deceptive information. In this paper, we study a…

人工智能 · 计算机科学 2026-04-14 Keyang Zhong , Junlin Xie , Hefeng Wu , Haofeng Li , Guanbin Li

Text-to-video (T2V) generation has rapidly progressed in visual fidelity, yet its ability to faithfully represent multiple cultures within a single prompt remains underexplored. We introduce MAVEN, a multi-agent prompt refinement framework…

计算机视觉与模式识别 · 计算机科学 2026-05-28 Shuowei Li , Yuming Zhao , Parth Bhalerao , Oana Ignat

Creating data stories from raw data is challenging due to humans' limited attention spans and the need for specialized skills. Recent advancements in large language models (LLMs) offer great opportunities to develop systems with autonomous…

人机交互 · 计算机科学 2024-08-08 Leixian Shen , Haotian Li , Yun Wang , Huamin Qu

Long-horizon omnimodal question answering answers questions by reasoning over text, images, audio, and video. Despite recent progress on OmniLLMs, low-resource long audio-video QA still suffers from costly dense encoding, weak fine-grained…

计算与语言 · 计算机科学 2026-03-31 Yifan Zhu , Xinyu Mu , Tao Feng , Zhonghong Ou , Yuning Gong , Haoran Luo

MLLMs have been widely studied for video question answering recently. However, most existing assessments focus on natural videos, overlooking synthetic videos, such as AI-generated content (AIGC). Meanwhile, some works in video generation…

计算机视觉与模式识别 · 计算机科学 2025-05-30 Tingyu Song , Tongyan Hu , Guo Gan , Yilun Zhao
‹ 上一页 1 2 3 10 下一页 ›