English
Related papers

Related papers: RoboEgo System Card: An Omnimodal Model with Nativ…

200 papers

Real-world live retrieval-augmented generation (RAG) systems face significant challenges when processing user queries that are often noisy, ambiguous, and contain multiple intents. While RAG enhances large language models (LLMs) with…

Computation and Language · Computer Science 2025-06-27 Guanting Dong , Xiaoxi Li , Yuyao Zhang , Mengjie Deng

Grounding textual expressions on scene objects from first-person views is a truly demanding capability in developing agents that are aware of their surroundings and behave following intuitive text instructions. Such capability is of…

Computer Vision and Pattern Recognition · Computer Science 2023-10-31 Shuhei Kurita , Naoki Katsura , Eri Onami

Robots must integrate multiple sensory modalities to act effectively in the real world. Yet, learning such multimodal policies at scale remains challenging. Simulation offers a viable solution, but while vision has benefited from…

Currently, medical vision language models are widely used in medical vision question answering tasks. However, existing models are confronted with two issues: for input, the model only relies on text instructions and lacks direct…

Computer Vision and Pattern Recognition · Computer Science 2025-10-14 Yanyuan Chen , Dexuan Xu , Yu Huang , Songkun Zhan , Hanpin Wang , Dongxue Chen , Xueping Wang , Meikang Qiu , Hang Li

Recently, humanoid robots have made significant advances in their ability to perform challenging tasks due to the deployment of Reinforcement Learning (RL), however, the inherent complexity of humanoid robots, including the difficulty of…

In today's world, emotional support is increasingly essential, yet it remains challenging for both those seeking help and those offering it. Multimodal approaches to emotional support show great promise by integrating diverse data sources…

Recently, human-computer interaction with various modalities has shown promising applications, like GPT-4o and Gemini. Given the foundational role of multimodal joint representation in understanding and generation pipelines, high-quality…

Computer Vision and Pattern Recognition · Computer Science 2024-07-17 Zehan Wang , Ziang Zhang , Hang Zhang , Luping Liu , Rongjie Huang , Xize Cheng , Hengshuang Zhao , Zhou Zhao

Artificial Intelligence (AI) has significantly advanced in recent years, driving innovation across various fields, especially in robotics. Even though robots can perform complex tasks with increasing autonomy, challenges remain in ensuring…

Human-Computer Interaction · Computer Science 2025-03-24 Anargh Viswanath , Lokesh Veeramacheneni , Hendrik Buschmeier

World models are widely explored in embodied intelligence, yet they typically predict distinct evolutions of the world and the ego within a single stream, where the world captures persistent instruction-agnostic scene regularities and the…

Computer Vision and Pattern Recognition · Computer Science 2026-05-20 Zuyao Lin , Jianhui Zhang , Peidong Jia , Xiaoguang Zhao , Shanghang Zhang , Xingyu Chen

Recent advances in large language models (LLMs) have driven impressive progress in omni-modal understanding and generation. However, training omni-modal LLMs remains a significant challenge due to the heterogeneous model architectures…

Computation and Language · Computer Science 2025-08-08 Qianli Ma , Yaowei Zheng , Zhelun Shi , Zhongkai Zhao , Bin Jia , Ziyue Huang , Zhiqi Lin , Youjie Li , Jiacheng Yang , Yanghua Peng , Zhi Zhang , Xin Liu

Situated embodied conversation requires robots to interleave real-time dialogue with active perception: deciding what to look at, when to look, and what to say under tight latency constraints. We present a simple, minimal system recipe that…

Robotics · Computer Science 2026-02-05 Dong Won Lee , Sarah Gillet , Louis-Philippe Morency , Cynthia Breazeal , Hae Won Park

As autonomous driving technology matures, end-to-end methodologies have emerged as a leading strategy, promising seamless integration from perception to control via deep learning. However, existing systems grapple with challenges such as…

Integrating Large Language Models (VLMs) and Vision-Language Models (VLMs) with robotic systems enables robots to process and understand complex natural language instructions and visual information. However, a fundamental challenge remains:…

Robotics · Computer Science 2024-03-18 Yuhang Hu , Yunzhe Wang , Ruibo Liu , Zhou Shen , Hod Lipson

Emotions conveyed through voice and face shape engagement and context in human AI interaction. Despite rapid progress in omni modal large language models, the holistic evaluation of emotional reasoning with audiovisual cues remains limited.…

Edge robotics involves frequent exchanges of large-volume multi-modal data. Existing methods ignore the interdependency between robotic functionalities and communication conditions, leading to excessive communication overhead. This paper…

Robotics · Computer Science 2025-10-21 Dan Guo , Xibin Jin , Shuai Wang , Zhigang Wen , Miaowen Wen , Chengzhong Xu

It is crucial that robots' performance can be improved after deployment, as they are inherently likely to encounter novel scenarios never seen before. This paper presents an innovative solution: an interactive learning-based robot system…

Human-Computer Interaction · Computer Science 2025-08-01 Kohou Wang , ZhaoXiang Liu , Lin Bai , Kun Fan , Xiang Liu , Huan Hu , Kai Wang , Shiguo Lian

When designing robots to assist in everyday human activities, it is crucial to enhance user requests with visual cues from their surroundings for improved intent understanding. This process is defined as a multimodal classification task.…

Computation and Language · Computer Science 2025-06-18 Shang-Chi Tsai , Seiya Kawano , Angel Garcia Contreras , Koichiro Yoshino , Yun-Nung Chen

While data-driven imitation learning has revolutionized robotic manipulation, current approaches remain constrained by the scarcity of large-scale, diverse real-world demonstrations. Consequently, the ability of existing models to…

The integration of electric vehicles (EVs) into smart grids presents unique opportunities to enhance both transportation systems and energy networks. However, ensuring safe and interpretable interactions between drivers, vehicles, and the…

Artificial Intelligence · Computer Science 2025-10-06 Jean Douglas Carvalho , Hugo Kenji , Ahmad Mohammad Saber , Glaucia Melo , Max Mauro Dias Santos , Deepa Kundur

The ability to navigate and interact with complex environments is central to real-world embodied agents, yet navigation in unseen environments remains challenging due to "experiential amnesia," where existing trajectory-driven or reactive…