English
Related papers

Related papers: X-Capture: An Open-Source Portable Device for Mult…

200 papers

Transparent objects are common in our daily life and frequently handled in the automated production line. Robust vision-based robotic grasping and manipulation for these objects would be beneficial for automation. However, the majority of…

Robotics · Computer Science 2022-08-30 Hongjie Fang , Hao-Shu Fang , Sheng Xu , Cewu Lu

From scientific research to commercial applications, eye tracking is an important tool across many domains. Despite its range of applications, eye tracking has yet to become a pervasive technology. We believe that we can put the power of…

Computer Vision and Pattern Recognition · Computer Science 2016-06-21 Kyle Krafka , Aditya Khosla , Petr Kellnhofer , Harini Kannan , Suchendra Bhandarkar , Wojciech Matusik , Antonio Torralba

The multitude of data generated by sensors available on users' mobile devices, combined with advances in machine learning techniques, support context-aware services in recognizing the current situation of a user (i.e., physical context) and…

Machine Learning · Computer Science 2023-06-29 Mattia Giovanni Campana , Dimitris Chatzopoulos , Franca Delmastro , Pan Hui

Learning effective multi-modal 3D representations of objects is essential for numerous applications, such as augmented reality and robotics. Existing methods often rely on task-specific embeddings that are tailored either for semantic…

Computer Vision and Pattern Recognition · Computer Science 2025-11-06 Gaia Di Lorenzo , Federico Tombari , Marc Pollefeys , Daniel Barath

Humans use all of their senses to accomplish different tasks in everyday activities. In contrast, existing work on robotic manipulation mostly relies on one, or occasionally two modalities, such as vision and touch. In this work, we…

Reliable embodied perception from an egocentric perspective is challenging yet essential for autonomous navigation technology of intelligent mobile agents. With the growing demand of social robotics, near-field scene understanding becomes…

Computer Vision and Pattern Recognition · Computer Science 2025-03-06 Haisheng Su , Feixiang Song , Cong Ma , Wei Wu , Junchi Yan

Building multisensory AI systems that learn from multiple sensory inputs such as text, speech, video, real-world sensors, wearable devices, and medical data holds great promise for impact in many scientific areas with practical benefits,…

Machine Learning · Computer Science 2024-05-01 Paul Pu Liang

Tactile-aware robot learning faces critical challenges in data collection and representation due to data scarcity and sparsity, and the absence of force feedback in existing systems. To address these limitations, we introduce a tactile…

Robotics · Computer Science 2025-09-19 Yue Xu , Litao Wei , Pengyu An , Qingyu Zhang , Yong-Lu Li

Object recognition and object pose estimation in robotic grasping continue to be significant challenges, since building a labelled dataset can be time consuming and financially costly in terms of data collection and annotation. In this…

Computer Vision and Pattern Recognition · Computer Science 2024-01-25 Dongmyoung Lee , Wei Chen , Nicolas Rojas

Cross-modal transfer learning is used to improve multi-modal classification models (e.g., for human activity recognition in human-robot collaboration). However, existing methods require paired sensor data at both training and inference,…

Machine Learning · Computer Science 2025-09-15 Leen Daher , Zhaobo Wang , Malcolm Mielle

Learning meaningful and compact representations with disentangled semantic aspects is considered to be of key importance in representation learning. Since real-world data is notoriously costly to collect, many recent state-of-the-art…

We tackle a challenging task: multi-view and multi-modal event detection that detects events in a wide-range real environment by utilizing data from distributed cameras and microphones and their weak labels. In this task, distributed…

Audio and Speech Processing · Electrical Eng. & Systems 2022-02-21 Masahiro Yasuda , Yasunori Ohishi , Shoichiro Saito , Noboru Harada

Developing robust multi-modal feature representations is crucial for enhancing object tracking performance. In pursuit of this objective, a novel X Modality Assisting Network (X-Net) is introduced, which explores the impact of the fusion…

Computer Vision and Pattern Recognition · Computer Science 2025-02-25 Zhaisheng Ding , Haiyan Li , Ruichao Hou , Yanyu Liu , Shidong Xie

We introduce a data capture system and a new dataset, HO-Cap, for 3D reconstruction and pose tracking of hands and objects in videos. The system leverages multiple RGBD cameras and a HoloLens headset for data collection, avoiding the use of…

Computer Vision and Pattern Recognition · Computer Science 2025-03-12 Jikai Wang , Qifan Zhang , Yu-Wei Chao , Bowen Wen , Xiaohu Guo , Yu Xiang

As humans, we experience the world with all our senses or modalities (sound, sight, touch, smell, and taste). We use these modalities, particularly sight and touch, to convey and interpret specific meanings. Multimodal expressions are…

Machine Learning · Computer Science 2022-05-17 Anirudh Sundar , Larry Heck

While current personal smart devices excel in digital domains, they fall short in assisting users during human environment interaction. This paper proposes Heads Up eXperience (HUX), an AI system designed to bridge this gap, serving as a…

Human-Computer Interaction · Computer Science 2024-07-30 Sukanth K , Sudhiksha Kandavel Rajan , Rajashekhar V S , Gowdham Prabhakar

Multimodal object recognition is still an emerging field. Thus, publicly available datasets are still rare and of small size. This dataset was developed to help fill this void and presents multimodal data for 63 objects with some visual and…

This paper introduces OmniMotion-X, a versatile multimodal framework for whole-body human motion generation, leveraging an autoregressive diffusion transformer in a unified sequence-to-sequence manner. OmniMotion-X efficiently supports…

Computer Vision and Pattern Recognition · Computer Science 2025-10-23 Guowei Xu , Yuxuan Bian , Ailing Zeng , Mingyi Shi , Shaoli Huang , Wen Li , Lixin Duan , Qiang Xu

The most common sensing modalities found in a robot perception system are vision and touch, which together can provide global and highly localized data for manipulation. However, these sensing modalities often fail to adequately capture the…

Robotics · Computer Science 2022-04-20 Jessica Yin , Gregory M. Campbell , James Pikul , Mark Yim

Humans naturally interact with both others and the surrounding multiple objects, engaging in various social activities. However, recent advances in modeling human-object interactions mostly focus on perceiving isolated individuals and…

Computer Vision and Pattern Recognition · Computer Science 2024-04-03 Juze Zhang , Jingyan Zhang , Zining Song , Zhanhe Shi , Chengfeng Zhao , Ye Shi , Jingyi Yu , Lan Xu , Jingya Wang