English
Related papers

Related papers: See-Control: A Multimodal Agent Framework for Smar…

200 papers

The human ability to seamlessly perform multimodal reasoning and physical interaction in the open world is a core goal for general purpose embodied intelligent systems. Recent vision-language-action (VLA) models, which are co-trained on…

Embodied AI systems, including AI-powered robots that autonomously interact with the physical world, stand to be significantly advanced by Large Language Models (LLMs), which enable robots to better understand complex language commands and…

Robotics · Computer Science 2024-09-04 Wenxiao Zhang , Xiangrui Kong , Thomas Braunl , Jin B. Hong

In embodied AI, visual perception should be active rather than passive: the system must decide where to look and at what scale to sense to acquire maximally informative data under pixel and spatial budget constraints. Existing vision models…

Robotics · Computer Science 2026-04-06 Jiashu Yang , Yifan Han , Yucheng Xie , Ning Guo , Wenzhao Lian

Wearable robotic hand rehabilitation devices can allow greater freedom and flexibility than their workstation-like counterparts. However, the field is generally lacking effective methods by which the user can operate the device: such…

Robotics · Computer Science 2019-01-14 Sangwoo Park , Cassie Meeker , Lynne M. Weber , Lauri Bishop , Joel Stein , Matei Ciocarlie

Effective collaboration between embodied agents requires more than acting in a shared environment; it demands communication grounded in each agent's evolving understanding of the world. When agents can only partially observe their…

Multiagent Systems · Computer Science 2026-05-19 Vardhan Dongre , Dilek Hakkani-Tür

We propose a novel hands-free control framework for the Boston Dynamics Spot robot using the Microsoft HoloLens 2 mixed-reality headset. Enabling accessible robot control is critical for allowing individuals with physical disabilities to…

Visual assistive technologies, such as Microsoft Seeing AI, can improve access to environmental information for persons with blindness or low vision (pBLV). Yet, the physical and functional implications of different device embodiments…

Human-Computer Interaction · Computer Science 2026-04-30 Gaurav Seth , Hoa Pham , Giles Hamilton-Fletcher , Charles Leclercq , John-Ross Rizzo

Recent advances in control robot methods, from end-to-end vision-language-action frameworks to modular systems with predefined primitives, have advanced robots' ability to follow natural language instructions. Nonetheless, many approaches…

Heterogeneous multi-robot systems (HMRS) have emerged as a powerful approach for tackling complex tasks that single robots cannot manage alone. Current large-language-model-based multi-agent systems (LLM-based MAS) have shown success in…

Robotics · Computer Science 2025-02-18 Junting Chen , Checheng Yu , Xunzhe Zhou , Tianqi Xu , Yao Mu , Mengkang Hu , Wenqi Shao , Yikai Wang , Guohao Li , Lin Shao

In Android GUI testing, generating an action sequence for a task that can be replayed as a test script is common. Generating sequences of actions and respective test scripts from task goals described in natural language can eliminate the…

Software Engineering · Computer Science 2025-09-12 Hieu Huynh , Hai Phung , Hao Pham , Tien N. Nguyen , Vu Nguyen

Smartphones have become indispensable in modern life, yet navigating complex tasks on mobile devices often remains frustrating. Recent advancements in large multimodal model (LMM)-based mobile agents have demonstrated the ability to…

Computation and Language · Computer Science 2025-01-29 Zhenhailong Wang , Haiyang Xu , Junyang Wang , Xi Zhang , Ming Yan , Ji Zhang , Fei Huang , Heng Ji

When assisting people in daily tasks, robots need to accurately interpret visual cues and respond effectively in diverse safety-critical situations, such as sharp objects on the floor. In this context, we present M-CoDAL, a…

Robotics · Computer Science 2025-02-26 Sabit Hassan , Hye-Young Chung , Xiang Zhi Tan , Malihe Alikhani

In this work, we present a multimodal system for active robot-object interaction using laser-based SLAM, RGBD images, and contact sensors. In the object manipulation task, the robot adjusts its initial pose with respect to obstacles and…

Robotics · Computer Science 2018-09-11 Luis Contreras , Hiroki Yokoyama , Hiroyuki Okada

Autonomous Earth Observation (EO) agents are transitioning from passive perception to complex, multi-step task execution. However, current architectures that integrate planning and execution within a single model often struggle with…

This paper describes, how current Machine Learning (ML) techniques combined with simple rule-based animation routines make an android robot head an embodied conversational agent with ChatGPT as its core component. The android robot head is…

Robotics · Computer Science 2024-01-11 Marcel Heisler , Christian Becker-Asano

Recent advancements in Large Language Models (LLMs) have greatly enhanced natural language understanding and content generation. However, these models primarily operate in disembodied digital environments and lack interaction with the…

Systems and Control · Electrical Eng. & Systems 2025-10-21 Wenbing Tang , Meilin Zhu , Fenghua Wu , Yang Liu

This paper introduces a novel mobile phone control architecture, Lightweight Multi-modal App Control (LiMAC), for efficient interactions and control across various Android apps. LiMAC takes as input a textual goal and a sequence of past…

Artificial Intelligence · Computer Science 2025-02-13 Filippos Christianos , Georgios Papoudakis , Thomas Coste , Jianye Hao , Jun Wang , Kun Shao

In embodied artificial intelligence, enabling heterogeneous robot teams to execute long-horizon tasks from high-level instructions remains a critical challenge. While large language models (LLMs) show promise in instruction parsing and…

Robotics · Computer Science 2026-03-06 Haishan Zeng , Mengna Wang , Peng Li

The growing dependence on mobile phones and their apps has made multi-user interactive features, like chat calls, live streaming, and video conferencing, indispensable for bridging the gaps in social connectivity caused by physical and…

Software Engineering · Computer Science 2025-09-17 Sidong Feng , Changhao Du , Huaxiao Liu , Qingnan Wang , Zhengwei Lv , Mengfei Wang , Chunyang Chen

Leveraging Multi-modal Large Language Models (MLLMs) to create embodied agents offers a promising avenue for tackling real-world tasks. While language-centric embodied agents have garnered substantial attention, MLLM-based embodied agents…