English
Related papers

Related papers: AIRoA MoMa Dataset: A Large-Scale Hierarchical Dat…

200 papers

Vision-Language MOT is a crucial tracking problem and has drawn increasing attention recently. It aims to track objects based on human language commands, replacing the traditional use of templates or pre-set information from training sets…

Computer Vision and Pattern Recognition · Computer Science 2024-06-13 Yunhao Li , Xiaoqiong Liu , Luke Liu , Heng Fan , Libo Zhang

Large-scale high-quality 3D motion datasets with multi-person interactions are crucial for data-driven models in autonomous driving to achieve fine-grained pedestrian interaction understanding in dynamic urban environments. However,…

Computer Vision and Pattern Recognition · Computer Science 2025-08-14 Guangxun Zhu , Shiyu Fan , Hang Dai , Edmond S. L. Ho

This work presents MAD (Multimodal Affection Dataset), a multimodal emotion dataset designed for affective computing and neurophysiological modeling. MAD is built upon synchronous collection of diverse physiological signals (EEG, ECG, EOG,…

Signal Processing · Electrical Eng. & Systems 2026-03-09 Shengwei Guo , Yunqing Qiao , Wenzhan Zhang , Bo Liu , Yong Wang , Guobing Sun

The CREATE database is composed of 14 hours of multimodal recordings from a mobile robotic platform based on the iRobot Create. The various sensors cover vision, audition, motors and proprioception. The dataset has been designed in the…

Robotics · Computer Science 2018-02-01 Simon Brodeur , Simon Carrier , Jean Rouat

Translating human intent into robot commands is crucial for the future of service robots in an aging society. Existing Human-Robot Interaction (HRI) systems relying on gestures or verbal commands are impractical for the elderly due to…

We present Galaxea Open-World Dataset, a large-scale, diverse collection of robot behaviors recorded in authentic human living and working environments. All demonstrations are gathered using a consistent robotic embodiment, paired with…

Robotics · Computer Science 2025-09-03 Tao Jiang , Tianyuan Yuan , Yicheng Liu , Chenhao Lu , Jianning Cui , Xiao Liu , Shuiqi Cheng , Jiyang Gao , Huazhe Xu , Hang Zhao

We used a 3D simulator to create artificial video data with standardized annotations, aiming to aid in the development of Embodied AI. Our question answering (QA) dataset measures the extent to which a robot can understand human behavior…

Artificial Intelligence · Computer Science 2024-09-18 Takanori Ugai , Kensho Hara , Shusaku Egami , Ken Fukuda

Human-machine interaction has been around for several decades now, with new applications emerging every day. One of the major goals that remain to be achieved is designing an interaction similar to how a human interacts with another human.…

Human-Computer Interaction · Computer Science 2022-12-27 Tauheed Khan Mohd , Nicole Nguyen , Ahmad Y Javaid

Time-series data are critical in diverse applications, such as industrial monitoring, medical diagnostics, and climate research. However, effectively integrating these high-dimensional temporal signals with natural language for dynamic,…

Computation and Language · Computer Science 2025-06-26 Yilin Wang , Peixuan Lei , Jie Song , Yuzhe Hao , Tao Chen , Yuxuan Zhang , Lei Jia , Yuanxiang Li , Zhongyu Wei

Multimodal systems have great potential to assist humans in procedural activities, where people follow instructions to achieve their goals. Despite diverse application scenarios, systems are typically evaluated on traditional classification…

Computation and Language · Computer Science 2025-11-05 Kimihiro Hasegawa , Wiradee Imrattanatrai , Zhi-Qi Cheng , Masaki Asada , Susan Holm , Yuran Wang , Ken Fukuda , Teruko Mitamura

Along with the increasing use of unmanned aerial vehicles (UAVs), large volumes of aerial videos have been produced. It is unrealistic for humans to screen such big data and understand their contents. Hence methodological research on the…

Computer Vision and Pattern Recognition · Computer Science 2020-06-26 Lichao Mou , Yuansheng Hua , Pu Jin , Xiao Xiang Zhu

Recently, Vision-Language Models (VLMs) have achieved remarkable progress in multimodal tasks, and multimodal instruction data serves as the foundation for enhancing VLM capabilities. Despite the availability of several open-source…

Generalist humanoid motion trackers have recently achieved strong simulation metrics by scaling data and training, yet often remain brittle on hardware during sustained teleoperation due to interface- and dynamics-induced errors. We present…

The integration of conversational agents into our daily lives has become increasingly common, yet many of these agents cannot engage in deep interactions with humans. Despite this, there is a noticeable shortage of datasets that capture…

Human-Computer Interaction · Computer Science 2025-03-19 Mohammed Althubyani , Zhijin Meng , Shengyuan Xie , Cha Seung , Imran Razzak , Eduardo B. Sandoval , Baki Kocaballi , Francisco Cruz

Multi-object tracking (MOT) in UAV-based video is challenging due to variations in viewpoint, low resolution, and the presence of small objects. While other research on MOT dedicated to aerial videos primarily focuses on the academic aspect…

Computer Vision and Pattern Recognition · Computer Science 2025-02-07 Nhat-Tan Do , Nhi Ngoc-Yen Nguyen , Dieu-Phuong Nguyen , Trong-Hop Do

We present MONET, a new multimodal dataset captured using a thermal camera mounted on a drone that flew over rural areas, and recorded human and vehicle activities. We captured MONET to study the problem of object localisation and behaviour…

Demand for air travel is rising, straining existing aviation infrastructure. In the US, more than 90% of airport control towers are understaffed, falling short of FAA and union standards. This, in part, has contributed to an uptick in…

Manipulation of large objects over long horizons (such as carts in a warehouse) is an essential skill for deployable robotic systems. Large objects require mobile manipulation which involves simultaneous manipulation, navigation, and…

Robotics · Computer Science 2024-10-10 Yajvan Ravan , Zhutian Yang , Tao Chen , Tomás Lozano-Pérez , Leslie Pack Kaelbling

This paper presents LLaMo (Large Language and Human Motion Assistant), a multimodal framework for human motion instruction tuning. In contrast to conventional instruction-tuning approaches that convert non-linguistic inputs, such as video…

Artificial Intelligence · Computer Science 2025-03-27 Lei Li , Sen Jia , Jianhao Wang , Zhongyu Jiang , Feng Zhou , Ju Dai , Tianfang Zhang , Zongkai Wu , Jenq-Neng Hwang