中文
相关论文

相关论文: JARVIS-1: Open-World Multi-task Agents with Memory…

200 篇论文

For embodied agents, navigation is an important ability but not an isolated goal. Agents are also expected to perform specific tasks after reaching the target location, such as picking up objects and assembling them into a particular…

计算与语言 · 计算机科学 2020-11-17 Hyounghun Kim , Abhay Zala , Graham Burri , Hao Tan , Mohit Bansal

Real-world multimodal applications often require any-to-any capabilities, enabling both understanding and generation across modalities including text, image, audio, and video. However, integrating the strengths of autoregressive language…

机器学习 · 计算机科学 2025-08-15 Jiulin Li , Ping Huang , Yexin Li , Shuo Chen , Juewen Hu , Ye Tian

LLaVA-Plus is a general-purpose multimodal assistant that expands the capabilities of large multimodal models. It maintains a skill repository of pre-trained vision and vision-language models and can activate relevant tools based on users'…

计算机视觉与模式识别 · 计算机科学 2023-11-10 Shilong Liu , Hao Cheng , Haotian Liu , Hao Zhang , Feng Li , Tianhe Ren , Xueyan Zou , Jianwei Yang , Hang Su , Jun Zhu , Lei Zhang , Jianfeng Gao , Chunyuan Li

Photo retouching has become integral to contemporary visual storytelling, enabling users to capture aesthetics and express creativity. While professional tools such as Adobe Lightroom offer powerful capabilities, they demand substantial…

计算机视觉与模式识别 · 计算机科学 2025-06-24 Yunlong Lin , Zixu Lin , Kunjie Lin , Jinbin Bai , Panwang Pan , Chenxin Li , Haoyu Chen , Zhongdao Wang , Xinghao Ding , Wenbo Li , Shuicheng Yan

Embodied agents powered by large language models (LLMs), such as Voyager, promise open-ended competence in worlds such as Minecraft. However, when powered by open-weight LLMs they still falter on elementary tasks after domain-specific…

人工智能 · 计算机科学 2025-12-17 Mircea Lică , Ojas Shirekar , Baptiste Colle , Chirag Raman

Large Language Models (LLMs) have the capacity of performing complex scheduling in a multi-agent system and can coordinate these agents into completing sophisticated tasks that require extensive collaboration. However, despite the…

We present a challenging benchmark for the Open WorLd VISual question answering (OWLViz) task. OWLViz presents concise, unambiguous queries that require integrating multiple capabilities, including visual understanding, web exploration, and…

机器学习 · 计算机科学 2025-07-31 Thuy Nguyen , Dang Nguyen , Hoang Nguyen , Thuan Luong , Long Hoang Dang , Viet Dac Lai

Aerial Visual Object Search (AVOS) tasks in urban environments require Unmanned Aerial Vehicles (UAVs) to autonomously search for and identify target objects using visual and textual cues without external guidance. Existing approaches…

计算机视觉与模式识别 · 计算机科学 2025-05-15 Yatai Ji , Zhengqiu Zhu , Yong Zhao , Beidan Liu , Chen Gao , Yihao Zhao , Sihang Qiu , Yue Hu , Quanjun Yin , Yong Li

The ability of Language Models (LMs) to understand natural language makes them a powerful tool for parsing human instructions into task plans for autonomous robots. Unlike traditional planning methods that rely on domain-specific knowledge…

Recent progress in multimodal reasoning has enabled agents that interpret imagery, connect it with language, and execute structured analytical tasks. Extending these capabilities to remote sensing remains challenging, as models must reason…

The Joint Automated Repository for Various Integrated Simulations (JARVIS) is a unified platform for multiscale, multimodal, forward, and inverse materials design. It integrates diverse theoretical and experimental approaches, including…

材料科学 · 物理学 2025-08-12 Kamal Choudhary

Large language model (LLM) based agents have shown great potential in following human instructions and automatically completing various tasks. To complete a task, the agent needs to decompose it into easily executed steps by planning.…

计算与语言 · 计算机科学 2025-06-02 Weihong Du , Wenrui Liao , Binyu Yan , Hongru Liang , Anthony G. Cohn , Wenqiang Lei

Vision-centric perception systems struggle with unpredictable and coupled weather degradations in the wild. Current solutions are often limited, as they either depend on specific degradation priors or suffer from significant domain gaps. To…

计算机视觉与模式识别 · 计算机科学 2025-04-08 Yunlong Lin , Zixu Lin , Haoyu Chen , Panwang Pan , Chenxin Li , Sixiang Chen , Yeying Jin , Wenbo Li , Xinghao Ding

Although numerous strategies have recently been proposed to enhance the autonomous interaction capabilities of multimodal agents in graphical user interface (GUI), their reliability remains limited when faced with complex or out-of-domain…

计算与语言 · 计算机科学 2025-10-06 Pengzhou Cheng , Lingzhong Dong , Zeng Wu , Zongru Wu , Xiangru Tang , Chengwei Qin , Zhuosheng Zhang , Gongshen Liu

Reinforcement learning agents must generalize beyond their training experience. Prior work has focused mostly on identical training and evaluation environments. Starting from the recently introduced Crafter benchmark, a 2D open world…

机器学习 · 计算机科学 2022-08-09 Aleksandar Stanić , Yujin Tang , David Ha , Jürgen Schmidhuber

Multimodal Large Language Models (MLLMs) have significantly advanced GUI agents, yet long-horizon automation remains constrained by two critical bottlenecks: context overload from raw sequential trajectory dependence and architectural…

人工智能 · 计算机科学 2026-04-15 Weihua Cheng , Junming Liu , Yifei Sun , Botian Shi , Yirong Chen , Ding Wang

Fine-grained understanding and species-specific multimodal question answering are vital for advancing biodiversity conservation and ecological monitoring. However, existing multimodal large language models face challenges when it comes to…

Large language models (LLMs) show remarkable potential to act as computer agents, enhancing human productivity and software accessibility in multi-modal tasks that require planning and reasoning. However, measuring agent performance in…

Video restoration in real-world scenarios is challenged by heterogeneous degradations, where static architectures and fixed inference pipelines often fail to generalize. Recent agent-based approaches offer dynamic decision making, yet…

计算机视觉与模式识别 · 计算机科学 2026-03-25 Xuanyu Zhang , Weiqi Li , Qunliang Xing , Jingfen Xie , Bin Chen , Junlin Li , Li Zhang , Jian Zhang , Shijie Zhao

A central challenge of visual control with model-based reinforcement learning (RL) is reliable long-horizon planning: long rollouts with learned latent dynamics exhibit branching futures and multi-modal action-value distributions. In…

机器学习 · 计算机科学 2026-05-07 Yurui Du , Pinhao Song , Yutong Hu , Renaud Detry