English
Related papers

Related papers: EO-Gym: A Multimodal, Interactive Environment for …

200 papers

Multi-hop audio-visual reasoning remains challenging for Omni-LLMs, as relevant evidence is often sparse, temporally dispersed, and distributed across both audio and visual streams. Existing benchmarks provide limited investigation of this…

Artificial Intelligence · Computer Science 2026-05-28 Ke Xu , Yuhao Wang , Ziyang Cheng , Hongcheng Liu , Yanfeng Wang , Yu Wang

We propose GAM-Agent, a game-theoretic multi-agent framework for enhancing vision-language reasoning. Unlike prior single-agent or monolithic models, GAM-Agent formulates the reasoning process as a non-zero-sum game between base…

Artificial Intelligence · Computer Science 2025-05-30 Jusheng Zhang , Yijia Fan , Wenjun Lin , Ruiqi Chen , Haoyi Jiang , Wenhao Chai , Jian Wang , Keze Wang

Both the design and control of a robot play equally important roles in its task performance. However, while optimal control is well studied in the machine learning and robotics community, less attention is placed on finding the optimal…

Robotics · Computer Science 2022-01-25 Jagdeep Singh Bhatia , Holly Jackson , Yunsheng Tian , Jie Xu , Wojciech Matusik

Vision-Language Models (VLMs) have advanced rapidly in multimodal perception and language understanding, yet it remains unclear whether they can reliably ground language into spatially coherent, plausibly executable actions in 3D digital…

Computer Vision and Pattern Recognition · Computer Science 2026-05-12 Niyati Rawal , Sushant Ravva , Shah Alam Abir , Saksham Jain , Aman Chadha , Vinija Jain , Suranjana Trivedy , Amitava Das

Autonomous UAV flight in confined, wall-dense environments requires low-latency and reliable motion planning under strict safety constraints. Traditional optimization-based planners suffer from mapping latency and easily fall into local…

Dynamic environments such as urban areas are still challenging for popular visual-inertial odometry (VIO) algorithms. Existing datasets typically fail to capture the dynamic nature of these environments, therefore making it difficult to…

Robotics · Computer Science 2021-02-12 Koji Minoda , Fabian Schilling , Valentin Wüest , Dario Floreano , Takehisa Yairi

Electromagnetismlike Optimization (EMO) is a global optimization algorithm, particularly well suited to solve problems featuring nonlinear and multimodal cost functions. EMO employs searcher agents that emulate a population of charged…

Artificial Intelligence · Computer Science 2014-05-21 Erik Cuevas , Diego Oliva , Daniel Zaldivar , Marco Perez , Gonzalo Pajares

Self-play has enabled large language models to autonomously improve through self-generated challenges. However, existing self-play methods for vision-language models rely on passive interaction with static image collections, resulting in…

Computer Vision and Pattern Recognition · Computer Science 2026-02-13 Jinghan He , Junfeng Fang , Feng Xiong , Zijun Yao , Fei Shen , Haiyun Guo , Jinqiao Wang , Tat-Seng Chua

Recent advances in Multimodal Large Language Models (MLLMs) have enabled their use as intelligent agents for smartphone operation. However, existing methods depend on the Android Debug Bridge (ADB) for data transmission and action…

Artificial Intelligence · Computer Science 2025-12-10 Haoyu Zhao , Weizhong Ding , Yuhao Yang , Zheng Tian , Linyi Yang , Kun Shao , Jun Wang

Simultaneous Localization and Mapping (SLAM) is considered to be an essential capability for intelligent vehicles and mobile robots. However, most of the current lidar SLAM approaches are based on the assumption of a static environment.…

Robotics · Computer Science 2022-06-22 Chenglong Qian , Zhaohong Xiang , Zhuoran Wu , Hongbin Sun

Visual Odometry (VO) is one of the fundamental tasks in computer vision for robotics. However, its performance is deeply affected by High Dynamic Range (HDR) scenes, omnipresent outdoor. While new Automatic-Exposure (AE) approaches to…

Learning a general motion tracking policy from human motions shows great potential for versatile humanoid whole-body control. Conventional approaches are not only inefficient in data utilization and training processes but also exhibit…

Robotics · Computer Science 2025-12-23 Chao Yang , Yingkai Sun , Peng Ye , Xin Chen , Chong Yu , Tao Chen

One of the most fundamental and information-laden actions humans do is to look at objects. However, a survey of current works reveals that existing gaze-related datasets annotate only the pixel being looked at, and not the boundaries of a…

Computer Vision and Pattern Recognition · Computer Science 2021-06-23 Henri Tomas , Marcus Reyes , Raimarc Dionido , Mark Ty , Jonric Mirando , Joel Casimiro , Rowel Atienza , Richard Guinto

Graphical User Interface (GUI) task automation constitutes a critical frontier in artificial intelligence research. While effective GUI agents synergistically integrate planning and grounding capabilities, current methodologies exhibit two…

Artificial Intelligence · Computer Science 2025-11-17 Yuan Zhao , Hualei Zhu , Tingyu Jiang , Shen Li , Xiaohang Xu , Hao Henry Wang

We present Coopetition-Gym v1, a benchmark platform for mixed-motive multi-agent reinforcement learning under strategic coopetition. The platform comprises twenty environments organized into four mechanism classes that correspond to four…

Multiagent Systems · Computer Science 2026-05-05 Vik Pant , Eric Yu

This paper addresses the problem of both actively searching and tracking multiple unknown dynamic objects in a known environment with multiple cooperative autonomous agents with partial observability. The tracking of a target ends when the…

As wearable and mobile devices become increasingly embedded in daily life, they offer a practical way to continuously sense human motion in the wild. But inertial signals are highly dependent on the sensing setup, including body location,…

Computer Vision and Pattern Recognition · Computer Science 2026-05-26 Baiyu Chen , Zechen Li , Wilson Wongso , Lihuan Li , Xiachong Lin , Hao Xue , Benjamin Tag , Flora Salim

This work presents SSL4EO-S12 v1.1, a multimodal, multitemporal Earth Observation dataset designed for pretraining large-scale foundation models. Building on the success of SSL4EO-S12, this extension updates the previous version to fix…

Computer Vision and Pattern Recognition · Computer Science 2026-02-18 Benedikt Blumenstiel , Nassim Ait Ali Braham , Conrad M Albrecht , Stefano Maurogiovanni , Paolo Fraccaro

Earth observation (EO) data such as satellite imagery can have far-reaching impacts on our understanding of the geography of poverty, especially when coupled with machine learning (ML) and computer vision. Early research used computer…

Machine Learning · Computer Science 2025-04-23 Kazuki Sakamoto , Connor T. Jerzak , Adel Daoud

We introduce WOFOSTGym, a novel crop simulation environment designed to train reinforcement learning (RL) agents to optimize agromanagement decisions for annual and perennial crops in single and multi-farm settings. Effective crop…

Artificial Intelligence · Computer Science 2025-02-28 William Solow , Sandhya Saisubramanian , Alan Fern
‹ Prev 1 8 9 10 Next ›