中文
相关论文

相关论文: UniVA: Universal Video Agent towards Open-Source N…

200 篇论文

Text-to-video generation has made significant strides, but replicating the capabilities of advanced systems like OpenAI Sora remains challenging due to their closed-source nature. Existing open-source methods struggle to achieve comparable…

计算机视觉与模式识别 · 计算机科学 2024-10-07 Zhengqing Yuan , Yixin Liu , Yihan Cao , Weixiang Sun , Haolong Jia , Ruoxi Chen , Zhaoxu Li , Bin Lin , Li Yuan , Lifang He , Chi Wang , Yanfang Ye , Lichao Sun

The integration of Large Language Models (LLMs) with specialized tools presents new opportunities for intelligent automation systems. However, orchestrating multiple LLM-driven agents to tackle complex tasks remains challenging due to…

人工智能 · 计算机科学 2025-03-27 Pengfei Du

Large Language Models (LLMs) have made the ambitious quest for generalist agents significantly far from being a fantasy. A key hurdle for building such general models is the diversity and heterogeneity of tasks and modalities. A promising…

计算机视觉与模式识别 · 计算机科学 2023-12-25 Mustafa Shukor , Corentin Dancette , Alexandre Rame , Matthieu Cord

Embodied navigation requires an agent to map language and visual observations to a stream of spatial actions that drive a real robot through environments it has never seen. The dominant approach has been to scale vision-language-action…

Unified understanding and generation is a highly appealing research direction in multimodal learning. There exist two approaches: one trains a transformer via an auto-regressive paradigm, and the other adopts a two-stage scheme connecting…

计算机视觉与模式识别 · 计算机科学 2025-12-09 Shihao Zhao , Yitong Chen , Zeyinzi Jiang , Bojia Zi , Shaozhe Hao , Yu Liu , Chaojie Mao , Kwan-Yee K. Wong

Text-to-video (T2V) generation has rapidly progressed in visual fidelity, yet its ability to faithfully represent multiple cultures within a single prompt remains underexplored. We introduce MAVEN, a multi-agent prompt refinement framework…

计算机视觉与模式识别 · 计算机科学 2026-05-28 Shuowei Li , Yuming Zhao , Parth Bhalerao , Oana Ignat

We present VINO, a unified visual generator that performs image and video generation and editing within a single framework. Instead of relying on task-specific models or independent modules for each modality, VINO uses a shared diffusion…

计算机视觉与模式识别 · 计算机科学 2026-01-19 Junyi Chen , Tong He , Zhoujie Fu , Pengfei Wan , Kun Gai , Weicai Ye

Unified multimodal models provide a natural and promising architecture for understanding diverse and complex real-world knowledge while generating high-quality images. However, they still rely primarily on frozen parametric knowledge, which…

With recent advances in multi-modal foundation models, the previously text-only large language models (LLM) have evolved to incorporate visual input, opening up unprecedented opportunities for various applications in visualization. Our work…

人机交互 · 计算机科学 2023-12-08 Shusen Liu , Haichao Miao , Zhimin Li , Matthew Olson , Valerio Pascucci , Peer-Timo Bremer

Image fusion aims to integrate complementary information from multiple source images to produce a more informative and visually consistent representation, benefiting both human perception and downstream vision tasks. Despite recent…

计算机视觉与模式识别 · 计算机科学 2026-03-17 Xingyuan Li , Songcheng Du , Yang Zou , HaoYuan Xu , Zhiying Jiang , Jinyuan Liu

Visual content creation tasks demand a nuanced understanding of design conventions and creative workflows-capabilities challenging for general models, while workflow-based agents lack specialized knowledge for autonomous creative planning.…

计算机视觉与模式识别 · 计算机科学 2026-03-04 Jinxiang Lai , Zexin Lu , Jiajun He , Rongwei Quan , Wenzhe Zhao , Qinyu Yang , Qi Chen , Qin Lin , Chuyue Li , Tao Gao , Yuhao Shan , Shuai Shao , Song Guo , Qinglin Lu

In this paper, we present JoVA, a unified framework for joint video-audio generation. Despite recent encouraging advances, existing methods face two critical limitations. First, most existing approaches can only generate ambient sounds and…

计算机视觉与模式识别 · 计算机科学 2025-12-16 Xiaohu Huang , Hao Zhou , Qiangpeng Yang , Shilei Wen , Kai Han

In recent years there have been remarkable breakthroughs in image-to-video generation. However, the 3D consistency and camera controllability of generated frames have remained unsolved. Recent studies have attempted to incorporate camera…

计算机视觉与模式识别 · 计算机科学 2024-10-15 Dejia Xu , Yifan Jiang , Chen Huang , Liangchen Song , Thorsten Gernoth , Liangliang Cao , Zhangyang Wang , Hao Tang

Generalization is a central challenge in autonomous driving, as real-world deployment requires robust performance under unseen scenarios, sensor domains, and environmental conditions. Recent world-model-based planning methods have shown…

计算机视觉与模式识别 · 计算机科学 2026-04-07 Mengmeng Liu , Diankun Zhang , Jiuming Liu , Jianfeng Cui , Hongwei Xie , Guang Chen , Hangjun Ye , Michael Ying Yang , Francesco Nex , Hao Cheng

The field of Abstract Visual Reasoning (AVR) encompasses a wide range of problems, many of which are inspired by human IQ tests. The variety of AVR tasks has resulted in state-of-the-art AVR methods being task-specific approaches.…

人工智能 · 计算机科学 2024-06-18 Mikołaj Małkiński , Jacek Mańdziuk

Medical diagnostic applications require models that can process multimodal medical inputs (images, patient histories, lab results) and generate diverse outputs including both textual reports and visual content (annotations, segmentation…

Large language models have demonstrated impressive universal capabilities across a wide range of open-ended tasks and have extended their utility to encompass multimodal conversations. However, existing methods encounter challenges in…

计算机视觉与模式识别 · 计算机科学 2024-04-08 Peng Jin , Ryuichi Takanobu , Wancai Zhang , Xiaochun Cao , Li Yuan

Document Visual Question Answering (DocVQA) remains challenging for existing Vision-Language Models (VLMs), especially under complex reasoning and multi-step workflows. Current approaches struggle to decompose intricate questions into…

计算机视觉与模式识别 · 计算机科学 2026-03-04 Aymen Lassoued , Mohamed Ali Souibgui , Yousri Kessentini

Recent advancements in multi-agent systems have demonstrated significant potential for enhancing creative task performance, such as long video generation. This study introduces three innovations to improve multi-agent collaboration. First,…

多智能体系统 · 计算机科学 2025-10-28 Zheng Wei , Mingchen Li , Zeqian Zhang , Ruibin Yuan , Pan Hui , Huamin Qu , James Evans , Maneesh Agrawala , Anyi Rao

Existing MLLMs encounter significant challenges in modeling the temporal context within long videos. Currently, mainstream Agent-based methods use external tools to assist a single MLLM in answering long video questions. Despite such…

计算机视觉与模式识别 · 计算机科学 2025-12-23 Boyu Chen , Zhengrong Yue , Siran Chen , Zikang Wang , Yang Liu , Peng Li , Yali Wang