English
Related papers

Related papers: UniMotion: Self-Supervised Learning for Cross-Doma…

200 papers

Touch is one of the most intuitive ways for humans to interact with the world, and as we advance toward a ubiquitous computing environment where technology seamlessly integrates into daily life, natural interaction methods are essential.…

Human-Computer Interaction · Computer Science 2024-12-24 Dev Shah

The image compression model has long struggled with adaptability and generalization, as the decoded bitstream typically serves only human or machine needs and fails to preserve information for unseen visual tasks. Therefore, this paper…

Computer Vision and Pattern Recognition · Computer Science 2025-01-09 Kangsheng Yin , Quan Liu , Xuelin Shen , Yulin He , Wenhan Yang , Shiqi Wang

Human perception of similarity across uni- and multimodal inputs is highly complex, making it challenging to develop automated metrics that accurately mimic it. General purpose vision-language models, such as CLIP and large multi-modal…

Computer Vision and Pattern Recognition · Computer Science 2024-12-17 Sara Ghazanfari , Siddharth Garg , Nicolas Flammarion , Prashanth Krishnamurthy , Farshad Khorrami , Francesco Croce

This paper proposes an approach for improving performance of unimodal models with multimodal training. Our approach involves a multi-branch architecture that incorporates unimodal models with a multimodal transformer-based branch. By…

Machine Learning · Computer Science 2023-11-20 Kateryna Chumachenko , Alexandros Iosifidis , Moncef Gabbouj

Gestures are inherent to human interaction and often complement speech in face-to-face communication, forming a multimodal communication system. An important task in gesture analysis is detecting a gesture's beginning and end. Research on…

Computer Vision and Pattern Recognition · Computer Science 2024-04-24 Esam Ghaleb , Ilya Burenko , Marlou Rasenberg , Wim Pouw , Ivan Toni , Peter Uhrig , Anna Wilson , Judith Holler , Aslı Özyürek , Raquel Fernández

Inertial Measurement Units (IMUs) are interceptive modalities that provide ego-motion measurements independent of the environmental factors. They are widely adopted in various autonomous systems. Motivated by the limitations in processing…

Machine Learning · Computer Science 2021-01-19 Rooholla Khorrambakht , Chris Xiaoxuan Lu , Hamed Damirchi , Zhenghua Chen , Zhengguo Li

Recent video generation models demonstrate impressive synthesis capabilities but remain limited by single-modality conditioning, constraining their holistic world understanding. This stems from insufficient cross-modal interaction and…

Computer Vision and Pattern Recognition · Computer Science 2025-12-09 Jiehui Huang , Yuechen Zhang , Xu He , Yuan Gao , Zhi Cen , Bin Xia , Yan Zhou , Xin Tao , Pengfei Wan , Jiaya Jia

Recent progress in large models has led to significant advances in unified multimodal generation and understanding. However, the development of models that unify motion-language generation and understanding remains largely underexplored.…

Computer Vision and Pattern Recognition · Computer Science 2026-04-20 Zekun Li , Sizhe An , Chengcheng Tang , Chuan Guo , Ivan Shugurov , Linguang Zhang , Amy Zhao , Srinath Sridhar , Lingling Tao , Abhay Mittal

Unified multimodal models (UMMs) aim to integrate understanding and generation within a single architecture. However, it remains underexplored how to effectively coordinate these two capabilities for more effective and efficient reasoning.…

Multimedia · Computer Science 2026-05-13 Hayes Bai , Yinyi Luo , Wenwen Wang , Qingsong Wen , Jindong Wang

Human gesture recognition using millimeter-wave (mmWave) signals provides attractive applications including smart home and in-car interfaces. While existing works achieve promising performance under controlled settings, practical…

Computer Vision and Pattern Recognition · Computer Science 2022-10-11 Yadong Li , Dongheng Zhang , Jinbo Chen , Jinwei Wan , Dong Zhang , Yang Hu , Qibin Sun , Yan Chen

We propose a deep video prediction model conditioned on a single image and an action class. To generate future frames, we first detect keypoints of a moving object and predict future motion as a sequence of keypoints. The input image is…

Computer Vision and Pattern Recognition · Computer Science 2019-10-07 Yunji Kim , Seonghyeon Nam , In Cho , Seon Joo Kim

We present UniBind, a flexible and efficient approach that learns a unified representation space for seven diverse modalities -- images, text, audio, point cloud, thermal, video, and event data. Existing works, eg., ImageBind, treat the…

Computer Vision and Pattern Recognition · Computer Science 2024-03-20 Yuanhuiyi Lyu , Xu Zheng , Jiazhou Zhou , Lin Wang

Combining different sensing modalities with multiple positions helps form a unified perception and understanding of complex situations such as human behavior. Hence, human activity recognition (HAR) benefits from combining redundant and…

Machine Learning · Computer Science 2024-04-26 Hymalai Bello

Recent advancements in millimeter-wave (mmWave) radar have demonstrated its potential for human action recognition and pose estimation, offering privacy-preserving advantages over conventional cameras while maintaining occlusion robustness,…

Computer Vision and Pattern Recognition · Computer Science 2025-07-08 Yizhe Lv , Tingting Zhang , Zhijian Wang , Yunpeng Song , Han Ding , Jinsong Han , Fei Wang

We consider the task of multimodal one-shot speech-image matching. An agent is shown a picture along with a spoken word describing the object in the picture, e.g. cookie, broccoli and ice-cream. After observing one paired speech-image…

Computation and Language · Computer Science 2020-08-17 Leanne Nortje , Herman Kamper

Recent research proposed eyelid gestures for people with upper-body motor impairments (UMI) to interact with smartphones without finger touch. However, such eyelid gestures were designed by researchers. It remains unknown what eyelid…

Human-Computer Interaction · Computer Science 2022-02-15 Xuan Zhao , Mingming Fan , Teng Han

Contact-rich manipulation requires reliable estimation of extrinsic contacts-the interactions between a grasped object and its environment which provide essential contextual information for planning, control, and policy learning. However,…

Robotics · Computer Science 2026-02-03 Zhengtong Xu , Yuki Shirai

Recent advances in Large Multi-modal Models (LMMs) have demonstrated their remarkable success as general-purpose multi-modal assistants, with particular focuses on holistic image- and video-language understanding. Conversely, less attention…

Computer Vision and Pattern Recognition · Computer Science 2025-11-11 Ye Liu , Zongyang Ma , Junfu Pu , Zhongang Qi , Yang Wu , Ying Shan , Chang Wen Chen

Human activity recognition (HAR) based on multimodal sensors has become a rapidly growing branch of biometric recognition and artificial intelligence. However, how to fully mine multimodal time series data and effectively learn accurate…

Computer Vision and Pattern Recognition · Computer Science 2022-05-25 Jialiang Wang , Haotian Wei , Yi Wang , Shu Yang , Chi Li

The recent success in human action recognition with deep learning methods mostly adopt the supervised learning paradigm, which requires significant amount of manually labeled data to achieve good performance. However, label collection is an…

Computer Vision and Pattern Recognition · Computer Science 2018-09-07 Junnan Li , Yongkang Wong , Qi Zhao , Mohan S. Kankanhalli