中文
相关论文

相关论文: AIRoA MoMa Dataset: A Large-Scale Hierarchical Dat…

200 篇论文

Humanoid robots capable of autonomous operation in diverse environments have long been a goal for roboticists. However, autonomous manipulation by humanoid robots has largely been restricted to one specific scene, primarily due to the…

机器人学 · 计算机科学 2025-09-10 Yanjie Ze , Zixuan Chen , Wenhao Wang , Tianyi Chen , Xialin He , Ying Yuan , Xue Bin Peng , Jiajun Wu

Existing datasets for 3D hand-object interaction are limited either in the data cardinality, data variations in interaction scenarios, or the quality of annotations. In this work, we present a comprehensive new training dataset for…

计算机视觉与模式识别 · 计算机科学 2024-09-09 Woojin Cho , Jihyun Lee , Minjae Yi , Minje Kim , Taeyun Woo , Donghwan Kim , Taewook Ha , Hyokeun Lee , Je-Hwan Ryu , Woontack Woo , Tae-Kyun Kim

In the field of multimodal large language models (MLLMs), common methods typically involve unfreezing the language model during training to foster profound visual understanding. However, the fine-tuning of such models with vision-language…

人工智能 · 计算机科学 2025-04-16 Bin Wang , Chunyu Xie , Dawei Leng , Yuhui Yin

Omnidirectional aerial robots offer full 6-DoF independent control over position and orientation, making them popular for aerial manipulation. Although advancements in robotic autonomy, human operation remains essential in complex aerial…

机器人学 · 计算机科学 2025-07-22 Jinjie Li , Jiaxuan Li , Kotaro Kaneko , Haokun Liu , Liming Shu , Moju Zhao

Scaling mobile manipulation imitation learning is bottlenecked by expensive mobile robot teleoperation. We present Egocentric Mobile MAnipulation (EMMA), an end-to-end framework training mobile manipulation policies from human mobile…

We present LiHRA, a novel dataset designed to facilitate the development of automated, learning-based, or classical risk monitoring (RM) methods for Human-Robot Interaction (HRI) scenarios. The growing prevalence of collaborative robots in…

机器人学 · 计算机科学 2025-12-01 Frederik Plahl , Georgios Katranis , Ilshat Mamaev , Andrey Morozov

We present HERO, a novel framework for large-scale video+language omni-representation learning. HERO encodes multimodal inputs in a hierarchical structure, where local context of a video frame is captured by a Cross-modal Transformer via…

计算机视觉与模式识别 · 计算机科学 2020-10-01 Linjie Li , Yen-Chun Chen , Yu Cheng , Zhe Gan , Licheng Yu , Jingjing Liu

Automatic analysis of teacher and student interactions could be very important to improve the quality of teaching and student engagement. However, despite some recent progress in utilizing multimodal data for teaching and learning…

计算机与社会 · 计算机科学 2022-12-07 Fangli Xu , Lingfei Wu , KP Thai , Carol Hsu , Wei Wang , Richard Tong

Validating Augmented Reality (AR) tracking and interaction models requires precise, repeatable ground-truth motion. However, human users cannot reliably perform consistent motion due to biomechanical variability. Robotic manipulators are…

机器人学 · 计算机科学 2026-02-09 Harsh Chhajed , Tian Guo

Significant progress has been made in open-vocabulary mobile manipulation, where the goal is for a robot to perform tasks in any environment given a natural language description. However, most current systems assume a static environment,…

Referring Multi-Object Tracking (RMOT) aims to achieve precise object detection and tracking through natural language instructions, representing a fundamental capability for intelligent robotic systems. However, current RMOT research…

Modelling interactions between humans and objects in natural environments is central to many applications including gaming, virtual and mixed reality, as well as human behavior analysis and human-robot collaboration. This challenging…

计算机视觉与模式识别 · 计算机科学 2022-04-15 Bharat Lal Bhatnagar , Xianghui Xie , Ilya A. Petrov , Cristian Sminchisescu , Christian Theobalt , Gerard Pons-Moll

Solving mobile manipulation tasks in inaccessible and dangerous environments is an important application of robots to support humans. Example domains are construction and maintenance of manned and unmanned stations on the moon and other…

In recent years, there have been significant advances in building end-to-end Machine Learning (ML) systems that learn at scale. But most of these systems are: (a) isolated (perception, speech, or language only); (b) trained on static…

In this paper, we propose the FoMo (For\^et Montmorency) dataset: a comprehensive, multi-season data collection. Located in the Montmorency Forest, Quebec, Canada, our dataset will capture a rich variety of sensory data over six distinct…

Robotic autonomy in open-world environments is fundamentally limited by insufficient data diversity and poor cross-embodiment generalization. Existing robotic datasets are often limited in scale and task coverage, while relatively large…

In this paper, we extended the method proposed in [21] to enable humans to interact naturally with autonomous agents through vocal and textual conversations. Our extended method exploits the inherent capabilities of pre-trained large…

机器人学 · 计算机科学 2024-12-31 Linus Nwankwo , Elmar Rueckert

Automated medical report generation for 3D PET/CT imaging is fundamentally challenged by the high-dimensional nature of volumetric data and a critical scarcity of annotated datasets, particularly for low-resource languages. Current…

Enterprise AI agents must continuously adapt to maintain accuracy, reduce latency, and remain aligned with user needs. We present a practical implementation of a data flywheel in NVInfo AI, NVIDIA's Mixture-of-Experts (MoE) Knowledge…

While robots deployed in real-world environments inevitably experience interaction failures, understanding how users respond through verbal and non-verbal behaviors remains under-explored in human-robot interaction (HRI). This gap is…