English
Related papers

Related papers: M2R2: MultiModal Robotic Representation for Tempor…

200 papers

Current video models fail as world model as they lack fine-graiend control. General-purpose household robots require real-time fine motor control to handle delicate tasks and urgent situations. In this work, we introduce fine-grained…

Computer Vision and Pattern Recognition · Computer Science 2025-10-03 Yichen Li , Antonio Torralba

Although the performance of Temporal Action Segmentation (TAS) has improved in recent years, achieving promising results often comes with a high computational cost due to dense inputs, complex model structures, and resource-intensive…

Computer Vision and Pattern Recognition · Computer Science 2024-05-28 Peiyao Wang , Yuewei Lin , Erik Blasch , Jie Wei , Haibin Ling

Incorporating additional sensory modalities such as tactile and audio into foundational robotic models poses significant challenges due to the curse of dimensionality. This work addresses this issue through modality selection. We propose a…

Robotics · Computer Science 2025-04-22 Jiawei Jiang , Kei Ota , Devesh K. Jha , Asako Kanezaki

Multimodal features play a key role in wearable sensor-based human activity recognition (HAR). Selecting the most salient features adaptively is a promising way to maximize the effectiveness of multimodal sensor data. In this regard, we…

Human-Computer Interaction · Computer Science 2018-05-21 Kaixuan Chen , Lina Yao , Xianzhi Wang , Dalin Zhang , Tao Gu , Zhiwen Yu , Zheng Yang

Understanding human actions from videos is essential in many domains, including sports. In figure skating, technical judgments are performed by watching skaters' 3D movements, and its part of the judging procedure can be regarded as a…

Computer Vision and Pattern Recognition · Computer Science 2024-08-30 Ryota Tanaka , Tomohiro Suzuki , Keisuke Fujii

Spatio-temporal feature learning is of central importance for action recognition in videos. Existing deep neural network models either learn spatial and temporal features independently (C2D) or jointly with unconstrained parameters (C3D).…

Computer Vision and Pattern Recognition · Computer Science 2019-03-05 Chao Li , Qiaoyong Zhong , Di Xie , Shiliang Pu

Multi-step manipulation tasks where robots interact with their environment and must apply process forces based on the perceived situation remain challenging to learn and prone to execution errors. Accurately simulating these tasks is also…

Robotics · Computer Science 2025-05-08 Christoph Willibald , Dongheui Lee

Foundation models pre-trained on massive unlabeled datasets have revolutionized natural language and computer vision, exhibiting remarkable generalization capabilities, thus highlighting the importance of pre-training. Yet, efforts in…

Robotics · Computer Science 2025-05-20 Dantong Niu , Yuvan Sharma , Haoru Xue , Giscard Biamby , Junyi Zhang , Ziteng Ji , Trevor Darrell , Roei Herzig

Human activity recognition (HAR) with wearables is promising research that can be widely adopted in many smart healthcare applications. In recent years, the deep learning-based HAR models have achieved impressive recognition performance.…

Computer Vision and Pattern Recognition · Computer Science 2022-08-17 Songpengcheng Xia , Lei Chu , Ling Pei , Wenxian Yu , Robert C. Qiu

Referring Image Segmentation (RIS) aims to segment target objects expressed in natural language within a scene at the pixel level. Various recent RIS models have achieved state-of-the-art performance by generating contextual tokens to model…

Computer Vision and Pattern Recognition · Computer Science 2023-12-01 Minhyeok Lee , Dogyoon Lee , Jungho Lee , Suhwan Cho , Heeseung Choi , Ig-Jae Kim , Sangyoun Lee

As of today, state-of-the-art activity recognition from wearable sensors relies on algorithms being trained to classify fixed windows of data. In contrast, video-based Human Activity Recognition, known as Temporal Action Localization (TAL),…

Machine Learning · Computer Science 2024-10-15 Marius Bock , Michael Moeller , Kristof Van Laerhoven

Dictionary learning algorithms have been successfully used for both reconstructive and discriminative tasks, where an input signal is represented with a sparse linear combination of dictionary atoms. While these methods are mostly developed…

Machine Learning · Statistics 2016-01-20 Soheil Bahrampour , Nasser M. Nasrabadi , Asok Ray , W. Kenneth Jenkins

Reasoning over dynamic visual content remains a central challenge for multimodal large language models. Recent thinking models generate explicit reasoning traces for interpretability; however, their reasoning often appears convincing while…

Computer Vision and Pattern Recognition · Computer Science 2025-12-09 Muhammad Maaz , Hanoona Rasheed , Fahad Shahbaz Khan , Salman Khan

A robot operating in unstructured environments must be able to discriminate between different grasping styles depending on the prospective manipulation task. Having a system that allows learning from remote non-expert demonstrations can…

Computer Vision and Pattern Recognition · Computer Science 2023-03-21 Furkan Kaynar , Sudarshan Rajagopalan , Shaobo Zhou , Eckehard Steinbach

This article presents MAPS$^2$ : a distributed algorithm that allows multi-robot systems to deliver coupled tasks expressed as Signal Temporal Logic (STL) constraints. Classical control theoretical tools addressing STL constraints either…

Robotics · Computer Science 2025-12-17 Mayank Sewlia , Christos K. Verginis , Dimos V. Dimarogonas

Temporal action localization has recently attracted significant interest in the Computer Vision community. However, despite the great progress, it is hard to identify which aspects of the proposed methods contribute most to the increase in…

Computer Vision and Pattern Recognition · Computer Science 2019-05-30 Tingting Xie , Xiaoshan Yang , Tianzhu Zhang , Changsheng Xu , Ioannis Patras

We study the task of robust feature representations, aiming to generalize well on multiple datasets for action recognition. We build our method on Transformers for its efficacy. Although we have witnessed great progress for video action…

Computer Vision and Pattern Recognition · Computer Science 2022-11-28 Junwei Liang , Enwei Zhang , Jun Zhang , Chunhua Shen

Deep learning models have enjoyed great success for image related computer vision tasks like image classification and object detection. For video related tasks like human action recognition, however, the advancements are not as significant…

Computer Vision and Pattern Recognition · Computer Science 2018-09-12 Xiaolin Song , Cuiling Lan , Wenjun Zeng , Junliang Xing , Jingyu Yang , Xiaoyan Sun

Promptable video object segmentation and tracking (VOST) has seen significant advances with the emergence of foundation models like Segment Anything Model 2 (SAM2); however, their application in surgical video analysis remains challenging…

Computer Vision and Pattern Recognition · Computer Science 2025-08-11 Guoping Xu , Hua-Chieh Shao , You Zhang

Robotic manipulation requires both rich multimodal perception and effective learning frameworks to handle complex real-world tasks. See-through-skin (STS) sensors, which combine tactile and visual perception, offer promising sensing…

Robotics · Computer Science 2026-02-10 Yuyang Li , Yinghan Chen , Zihang Zhao , Puhao Li , Tengyu Liu , Siyuan Huang , Yixin Zhu