中文
相关论文

相关论文: RoboEgo System Card: An Omnimodal Model with Nativ…

200 篇论文

Real-world live retrieval-augmented generation (RAG) systems face significant challenges when processing user queries that are often noisy, ambiguous, and contain multiple intents. While RAG enhances large language models (LLMs) with…

计算与语言 · 计算机科学 2025-06-27 Guanting Dong , Xiaoxi Li , Yuyao Zhang , Mengjie Deng

Grounding textual expressions on scene objects from first-person views is a truly demanding capability in developing agents that are aware of their surroundings and behave following intuitive text instructions. Such capability is of…

计算机视觉与模式识别 · 计算机科学 2023-10-31 Shuhei Kurita , Naoki Katsura , Eri Onami

Robots must integrate multiple sensory modalities to act effectively in the real world. Yet, learning such multimodal policies at scale remains challenging. Simulation offers a viable solution, but while vision has benefited from…

Currently, medical vision language models are widely used in medical vision question answering tasks. However, existing models are confronted with two issues: for input, the model only relies on text instructions and lacks direct…

计算机视觉与模式识别 · 计算机科学 2025-10-14 Yanyuan Chen , Dexuan Xu , Yu Huang , Songkun Zhan , Hanpin Wang , Dongxue Chen , Xueping Wang , Meikang Qiu , Hang Li

Recently, humanoid robots have made significant advances in their ability to perform challenging tasks due to the deployment of Reinforcement Learning (RL), however, the inherent complexity of humanoid robots, including the difficulty of…

机器人学 · 计算机科学 2024-08-27 Qiang Zhang , Peter Cui , David Yan , Jingkai Sun , Yiqun Duan , Gang Han , Wen Zhao , Weining Zhang , Yijie Guo , Arthur Zhang , Renjing Xu

In today's world, emotional support is increasingly essential, yet it remains challenging for both those seeking help and those offering it. Multimodal approaches to emotional support show great promise by integrating diverse data sources…

Recently, human-computer interaction with various modalities has shown promising applications, like GPT-4o and Gemini. Given the foundational role of multimodal joint representation in understanding and generation pipelines, high-quality…

计算机视觉与模式识别 · 计算机科学 2024-07-17 Zehan Wang , Ziang Zhang , Hang Zhang , Luping Liu , Rongjie Huang , Xize Cheng , Hengshuang Zhao , Zhou Zhao

Artificial Intelligence (AI) has significantly advanced in recent years, driving innovation across various fields, especially in robotics. Even though robots can perform complex tasks with increasing autonomy, challenges remain in ensuring…

人机交互 · 计算机科学 2025-03-24 Anargh Viswanath , Lokesh Veeramacheneni , Hendrik Buschmeier

World models are widely explored in embodied intelligence, yet they typically predict distinct evolutions of the world and the ego within a single stream, where the world captures persistent instruction-agnostic scene regularities and the…

计算机视觉与模式识别 · 计算机科学 2026-05-20 Zuyao Lin , Jianhui Zhang , Peidong Jia , Xiaoguang Zhao , Shanghang Zhang , Xingyu Chen

Recent advances in large language models (LLMs) have driven impressive progress in omni-modal understanding and generation. However, training omni-modal LLMs remains a significant challenge due to the heterogeneous model architectures…

计算与语言 · 计算机科学 2025-08-08 Qianli Ma , Yaowei Zheng , Zhelun Shi , Zhongkai Zhao , Bin Jia , Ziyue Huang , Zhiqi Lin , Youjie Li , Jiacheng Yang , Yanghua Peng , Zhi Zhang , Xin Liu

Situated embodied conversation requires robots to interleave real-time dialogue with active perception: deciding what to look at, when to look, and what to say under tight latency constraints. We present a simple, minimal system recipe that…

机器人学 · 计算机科学 2026-02-05 Dong Won Lee , Sarah Gillet , Louis-Philippe Morency , Cynthia Breazeal , Hae Won Park

As autonomous driving technology matures, end-to-end methodologies have emerged as a leading strategy, promising seamless integration from perception to control via deep learning. However, existing systems grapple with challenges such as…

机器人学 · 计算机科学 2023-10-27 Tsun-Hsuan Wang , Alaa Maalouf , Wei Xiao , Yutong Ban , Alexander Amini , Guy Rosman , Sertac Karaman , Daniela Rus

Integrating Large Language Models (VLMs) and Vision-Language Models (VLMs) with robotic systems enables robots to process and understand complex natural language instructions and visual information. However, a fundamental challenge remains:…

机器人学 · 计算机科学 2024-03-18 Yuhang Hu , Yunzhe Wang , Ruibo Liu , Zhou Shen , Hod Lipson

Emotions conveyed through voice and face shape engagement and context in human AI interaction. Despite rapid progress in omni modal large language models, the holistic evaluation of emotional reasoning with audiovisual cues remains limited.…

Edge robotics involves frequent exchanges of large-volume multi-modal data. Existing methods ignore the interdependency between robotic functionalities and communication conditions, leading to excessive communication overhead. This paper…

机器人学 · 计算机科学 2025-10-21 Dan Guo , Xibin Jin , Shuai Wang , Zhigang Wen , Miaowen Wen , Chengzhong Xu

It is crucial that robots' performance can be improved after deployment, as they are inherently likely to encounter novel scenarios never seen before. This paper presents an innovative solution: an interactive learning-based robot system…

人机交互 · 计算机科学 2025-08-01 Kohou Wang , ZhaoXiang Liu , Lin Bai , Kun Fan , Xiang Liu , Huan Hu , Kai Wang , Shiguo Lian

When designing robots to assist in everyday human activities, it is crucial to enhance user requests with visual cues from their surroundings for improved intent understanding. This process is defined as a multimodal classification task.…

计算与语言 · 计算机科学 2025-06-18 Shang-Chi Tsai , Seiya Kawano , Angel Garcia Contreras , Koichiro Yoshino , Yun-Nung Chen

While data-driven imitation learning has revolutionized robotic manipulation, current approaches remain constrained by the scarcity of large-scale, diverse real-world demonstrations. Consequently, the ability of existing models to…

The integration of electric vehicles (EVs) into smart grids presents unique opportunities to enhance both transportation systems and energy networks. However, ensuring safe and interpretable interactions between drivers, vehicles, and the…

The ability to navigate and interact with complex environments is central to real-world embodied agents, yet navigation in unseen environments remains challenging due to "experiential amnesia," where existing trajectory-driven or reactive…