中文
相关论文

相关论文: JARVIS-1: Open-World Multi-task Agents with Memory…

200 篇论文

Multimodal large language models (MLLMs) have shown remarkable capabilities in cross-modal understanding and reasoning, offering new opportunities for intelligent assistive systems, yet existing systems still struggle with risk-aware…

机器人学 · 计算机科学 2026-04-08 Renjun Gao

While Vision-Language Models (VLMs) hold promise for tasks requiring extensive collaboration, traditional multi-agent simulators have facilitated rich explorations of an interactive artificial society that reflects collective behavior.…

计算与语言 · 计算机科学 2024-05-24 Xianhao Yu , Jiaqi Fu , Renjia Deng , Wenjuan Han

Graphical User Interface (GUI) agents powered by Multimodal Large Language Models (MLLMs) promise human-like interaction with software applications, yet long-horizon tasks remain challenging due to memory limitations. Existing approaches…

计算机视觉与模式识别 · 计算机科学 2025-10-07 Zikang Liu , Junyi Li , Wayne Xin Zhao , Dawei Gao , Yaliang Li , Ji-rong Wen

Vision-language models (VLMs) have demonstrated remarkable capabilities in robotic planning, particularly for long-horizon tasks that require a holistic understanding of the environment for task decomposition. Existing methods typically…

机器人学 · 计算机科学 2025-03-31 Puzhen Yuan , Angyuan Ma , Yunchao Yao , Huaxiu Yao , Masayoshi Tomizuka , Mingyu Ding

The rapid progress of navigation, manipulation, and vision models has made mobile manipulators capable in many specialized tasks. However, the open-world mobile manipulation (OWMM) task remains a challenge due to the need for generalization…

机器人学 · 计算机科学 2025-06-24 Junting Chen , Haotian Liang , Lingxiao Du , Weiyun Wang , Mengkang Hu , Yao Mu , Wenhai Wang , Jifeng Dai , Ping Luo , Wenqi Shao , Lin Shao

The paper introduces GUI-Owl-1.5, the latest native GUI agent model that features instruct/thinking variants in multiple sizes (2B/4B/8B/32B/235B) and supports a range of platforms (desktop, mobile, browser, and more) to enable cloud-edge…

Remote sensing (RS) images from multiple modalities and platforms exhibit diverse details due to differences in sensor characteristics and imaging perspectives. Existing vision-language research in RS largely relies on relatively…

计算机视觉与模式识别 · 计算机科学 2026-01-06 Huiyang Hu , Peijin Wang , Yingchao Feng , Kaiwen Wei , Wenxin Yin , Wenhui Diao , Mengyu Wang , Hanbo Bi , Kaiyue Kang , Tong Ling , Kun Fu , Xian Sun

Many reinforcement learning environments (e.g., Minecraft) provide only sparse rewards that indicate task completion or failure with binary values. The challenge in exploration efficiency in such environments makes it difficult for…

人工智能 · 计算机科学 2024-04-02 Hao Li , Xue Yang , Zhaokai Wang , Xizhou Zhu , Jie Zhou , Yu Qiao , Xiaogang Wang , Hongsheng Li , Lewei Lu , Jifeng Dai

Real-world embodied agents face long-horizon tasks, characterized by high-level goals demanding multi-step solutions beyond single actions. Successfully navigating these requires both high-level task planning (i.e., decomposing goals into…

机器人学 · 计算机科学 2025-06-03 Yi Yang , Jiaxuan Sun , Siqi Kou , Yihan Wang , Zhijie Deng

Repurposing pre-trained diffusion models has been proven to be effective for NVS. However, these methods are mostly limited to a single object; directly applying such methods to compositional multi-object scenarios yields inferior results,…

计算机视觉与模式识别 · 计算机科学 2025-03-25 Ruijie Lu , Yixin Chen , Junfeng Ni , Baoxiong Jia , Yu Liu , Diwen Wan , Gang Zeng , Siyuan Huang

Completing Long-Horizon (LH) tasks in open-ended worlds is an important yet difficult problem for embodied agents. Existing approaches suffer from two key challenges: (1) they heavily rely on experiences obtained from human-created data or…

机器人学 · 计算机科学 2026-04-30 Tongtong Feng , Xin Wang , Zekai Zhou , Ren Wang , Yuwei Zhan , Guangyao Li , Qing Li , Wenwu Zhu

The rapid development of large language and multimodal models has sparked significant interest in using proprietary models, such as GPT-4o, to develop autonomous agents capable of handling real-world scenarios like web navigation. Although…

计算与语言 · 计算机科学 2024-10-28 Hongliang He , Wenlin Yao , Kaixin Ma , Wenhao Yu , Hongming Zhang , Tianqing Fang , Zhenzhong Lan , Dong Yu

A fundamental aspect of behaviour is the ability to encode salient features of experience in memory and use these memories, in combination with current sensory information, to predict the best action for each situation such that long-term…

神经与进化计算 · 计算机科学 2021-06-25 Stephen Kelly , Tatiana Voegerl , Wolfgang Banzhaf , Cedric Gondro

Video action detection (VAD) is a formidable vision task that involves the localization and classification of actions within the spatial and temporal dimensions of a video clip. Among the myriad VAD architectures, two-stage VAD methods…

计算机视觉与模式识别 · 计算机科学 2024-09-18 Seok Hwan Lee , Taein Son , Soo Won Seo , Jisong Kim , Jun Won Choi

Generalist models, which are capable of performing diverse multi-modal tasks in a task-agnostic way within a single model, have been explored recently. Being, hopefully, an alternative to approaching general-purpose AI, existing generalist…

计算机视觉与模式识别 · 计算机科学 2022-12-09 Jinze Bai , Rui Men , Hao Yang , Xuancheng Ren , Kai Dang , Yichang Zhang , Xiaohuan Zhou , Peng Wang , Sinan Tan , An Yang , Zeyu Cui , Yu Han , Shuai Bai , Wenbin Ge , Jianxin Ma , Junyang Lin , Jingren Zhou , Chang Zhou

Humanoid robots that autonomously interact with physical environments over extended horizons represent a central goal of embodied intelligence. Existing approaches rely on reference motions or task-specific rewards, tightly coupling…

机器人学 · 计算机科学 2026-02-26 Yutang Lin , Jieming Cui , Yixuan Li , Baoxiong Jia , Yixin Zhu , Siyuan Huang

As a fundamental problem for Artificial Intelligence, multi-agent system (MAS) is making rapid progress, mainly driven by multi-agent reinforcement learning (MARL) techniques. However, previous MARL methods largely focused on grid-world…

计算机视觉与模式识别 · 计算机科学 2021-07-21 Haiyang Wang , Wenguan Wang , Xizhou Zhu , Jifeng Dai , Liwei Wang

Recent advancements in Large Language Model~(LLM)-based Multi-Agent Systems (MAS) have demonstrated remarkable potential for tackling complex decision-making tasks. However, existing frameworks inevitably rely on serialized execution…

人工智能 · 计算机科学 2026-03-10 Yaoru Li , Shunyu Liu , Tongya Zheng , Li Sun , Mingli Song

We aim to develop a goal specification method that is semantically clear, spatially sensitive, domain-agnostic, and intuitive for human users to guide agent interactions in 3D environments. Specifically, we propose a novel cross-view goal…

人工智能 · 计算机科学 2025-07-10 Shaofei Cai , Zhancun Mu , Anji Liu , Yitao Liang

The quest for fully autonomous vehicles (AVs) capable of navigating complex real-world scenarios with human-like understanding and responsiveness. In this paper, we introduce Dolphins, a novel vision-language model architected to imbibe…

计算机视觉与模式识别 · 计算机科学 2023-12-04 Yingzi Ma , Yulong Cao , Jiachen Sun , Marco Pavone , Chaowei Xiao