中文
相关论文

相关论文: Unify-Agent: A Unified Multimodal Agent for World-…

200 篇论文

Unified models (UMs) hold promise for their ability to understand and generate content across heterogeneous modalities. Compared to merely generating visual content, the use of UMs for interleaved cross-modal reasoning is more promising and…

计算机视觉与模式识别 · 计算机科学 2026-04-03 Jiachun Jin , Zetong Zhou , Xiao Yang , Hao Zhang , Pengfei Liu , Jun Zhu , Zhijie Deng

AI agents today are mostly siloed - they either retrieve and reason over vast amount of digital information and knowledge obtained online; or interact with the physical world through embodied perception, planning and action - but rarely…

Millimeter-wave or terahertz communications can meet demands of low-altitude economy networks for high-throughput sensing and real-time decision making. However, high-frequency characteristics of wireless channels result in severe…

网络与互联网体系结构 · 计算机科学 2026-03-13 Min Hao , Zhizhuo Li , Zirui Zhang , Maoqiang Wu , Han Zhang , Rong Yu

We present PresentAgent, a multimodal agent that transforms long-form documents into narrated presentation videos. While existing approaches are limited to generating static slides or text summaries, our method advances beyond these…

计算机视觉与模式识别 · 计算机科学 2025-07-08 Jingwei Shi , Zeyu Zhang , Biao Wu , Yanjie Liang , Meng Fang , Ling Chen , Yang Zhao

Extracting actionable insights from complex value stream map simulations can be challenging, time-consuming, and error-prone. Recent advances in large language models offer new avenues to support users with this task. While existing…

计算与语言 · 计算机科学 2026-04-15 Micha Selak , Dirk Krechel , Adrian Ulges , Sven Spieckermann , Niklas Stoehr , Andreas Loehr

Despite advances in embodied AI, agent reasoning systems still struggle to capture the fundamental conceptual structures that humans naturally use to understand and interact with their environment. To address this, we propose a novel…

人工智能 · 计算机科学 2025-04-01 François Olivier , Zied Bouraoui

Omnimodal large language models have made significant strides in unifying audio and visual modalities; however, they often face challenges in fine-grained cross-modal understanding and have difficulty with multimodal alignment. To address…

计算机视觉与模式识别 · 计算机科学 2026-02-06 Keda Tao , Wenjie Du , Bohan Yu , Weiqiang Wang , Jian Liu , Huan Wang

Visual grounding is the task of localising image regions from natural language queries and is critical for reasoning capable Graphical User Interface agents. Many existing methods rely on massive, noisy synthetic datasets. This work…

人工智能 · 计算机科学 2025-11-17 Georgios Pantazopoulos , Eda B. Özyiğit

Pre-trained video models learn powerful priors for generating high-quality, temporally coherent content. While these models excel at temporal coherence, their dynamics are often constrained by the continuous nature of their training data.…

计算机视觉与模式识别 · 计算机科学 2025-12-02 Zhoujie Fu , Xianfang Zeng , Jinghong Lan , Xinyao Liao , Cheng Chen , Junyi Chen , Jiacheng Wei , Wei Cheng , Shiyu Liu , Yunuo Chen , Gang Yu , Guosheng Lin

Remote sensing (RS) images from multiple modalities and platforms exhibit diverse details due to differences in sensor characteristics and imaging perspectives. Existing vision-language research in RS largely relies on relatively…

计算机视觉与模式识别 · 计算机科学 2026-01-06 Huiyang Hu , Peijin Wang , Yingchao Feng , Kaiwen Wei , Wenxin Yin , Wenhui Diao , Mengyu Wang , Hanbo Bi , Kaiyue Kang , Tong Ling , Kun Fu , Xian Sun

Real-world industrial inspection requires not only localizing defects, but also explaining them in natural language and generating controlled defect edits. However, existing approaches fail to jointly support all three capabilities within a…

计算机视觉与模式识别 · 计算机科学 2026-04-15 Haoyu Zheng , Tianwei Lin , Wei Wang , Zhuonan Wang , Wenqiao Zhang , Jiaqi Zhu , Feifei Shao

Recent progress in using machine learning models for reasoning tasks has been driven by novel model architectures, large-scale pre-training protocols, and dedicated reasoning datasets for fine-tuning. In this work, to further pursue these…

机器学习 · 计算机科学 2023-09-18 Jack Lanchantin , Sainbayar Sukhbaatar , Gabriel Synnaeve , Yuxuan Sun , Kavya Srinet , Arthur Szlam

Recent significant advances in integrating multiple Large Language Model (LLM) systems have enabled Agentic Frameworks capable of performing complex tasks autonomously, including novel scientific research. We develop and demonstrate such a…

人工智能 · 计算机科学 2025-07-16 Darui Lu , Jordan M. Malof , Willie J. Padilla

While specialized AI models excel at isolated video tasks like generation or understanding, real-world applications demand complex, iterative workflows that combine these capabilities. To bridge this gap, we introduce UniVA, an open-source,…

计算机视觉与模式识别 · 计算机科学 2025-11-12 Zhengyang Liang , Daoan Zhang , Huichi Zhou , Rui Huang , Bobo Li , Yuechen Zhang , Shengqiong Wu , Xiaohan Wang , Jiebo Luo , Lizi Liao , Hao Fei

Image classification has traditionally relied on parameter-intensive model training, requiring large-scale annotated datasets and extensive fine tuning to achieve competitive performance. While recent vision language models (VLMs) alleviate…

计算机视觉与模式识别 · 计算机科学 2025-09-19 Wonduk Seo , Minhyeong Yu , Hyunjin An , Seunghyun Lee

This paper addresses the limitations of a single agent in task decomposition and collaboration during complex task execution, and proposes a multi-agent architecture for modular task decomposition and dynamic collaboration based on large…

人工智能 · 计算机科学 2025-11-04 Shuaidong Pan , Di Wu

The increasing realism of AI-Generated Images (AIGI) has created an urgent need for forensic tools capable of reliably distinguishing synthetic content from authentic imagery. Existing detectors are typically tailored to specific forgery…

计算机视觉与模式识别 · 计算机科学 2026-03-25 Yangxin Yu , Yue Zhou , Bin Li , Kaiqing Lin , Haodong Li , Jiangqun Ni , Bo Cao

Recent advancements in large language models (LLMs) have significantly improved the capabilities of web agents. However, effectively navigating complex and dynamic web environments still requires more advanced trajectory-level planning and…

人工智能 · 计算机科学 2025-07-08 Yifei Gao , Junhong Ye , Jiaqi Wang , Jitao Sang

Integrating image generation and understanding into a single framework has become a pivotal goal in the multimodal domain. However, how understanding can effectively assist generation has not been fully explored. Unlike previous works that…

计算机视觉与模式识别 · 计算机科学 2026-01-09 Yanbing Zeng , Jia Wang , Hanghang Ma , Junqiang Wu , Jie Zhu , Xiaoming Wei , Jie Hu

The next generation of autonomous agents must not only learn efficiently but also act reliably and adapt their behavior in open worlds. Standard approaches typically assume fixed tasks and environments with little or no novelty, which…

机器学习 · 计算机科学 2026-03-02 Florent Delgrange