English
Related papers

Related papers: ActionArt: Advancing Multimodal Large Models for F…

200 papers

The rapid progress of Multimodal Large Language Models (MLLMs) marks a significant step toward artificial general intelligence, offering great potential for augmenting human capabilities. However, their ability to provide effective…

Artificial Intelligence · Computer Science 2026-03-03 Hengjian Gao , Kaiwei Zhang , Shibo Wang , Mingjie Chen , Qihang Cao , Xianfeng Wang , Yucheng Zhu , Xiongkuo Min , Wei Sun , Dandan Zhu , Guangtao Zhai

Understanding bimanual human hand activities is a critical problem in AI and robotics. We cannot build large models of bimanual activities because existing datasets lack the scale, coverage of diverse hand activities, and detailed…

Computer Vision and Pattern Recognition · Computer Science 2025-04-10 Rao Fu , Dingxi Zhang , Alex Jiang , Wanjia Fu , Austin Funk , Daniel Ritchie , Srinath Sridhar

As mobile devices are becoming ubiquitous, regularly interacting with a variety of user interfaces (UIs) is a common aspect of daily life for many people. To improve the accessibility of these devices and to enable their usage in a variety…

Computation and Language · Computer Science 2021-01-27 Zecheng He , Srinivas Sunkara , Xiaoxue Zang , Ying Xu , Lijuan Liu , Nevan Wichers , Gabriel Schubiner , Ruby Lee , Jindong Chen , Blaise Agüera y Arcas

AI personal assistants, deployed through robots or wearables, require embodied understanding to collaborate effectively with humans. However, current Multimodal Large Language Models (MLLMs) primarily focus on third-person (exocentric)…

Computer Vision and Pattern Recognition · Computer Science 2025-12-17 Haoyu Zhang , Qiaohui Chu , Meng Liu , Haoxiang Shi , Yaowei Wang , Liqiang Nie

Wearable cameras allow to acquire images and videos from the user's perspective. These data can be processed to understand humans behavior. Despite human behavior analysis has been thoroughly investigated in third person vision, it is still…

Computer Vision and Pattern Recognition · Computer Science 2023-07-06 Francesco Ragusa , Antonino Furnari , Giovanni Maria Farinella

Action recognition is a vital task in computer vision, and many methods are developed to push it to the limit. However, current action recognition models have huge computational costs, which cannot be deployed to real-world tasks on mobile…

Computer Vision and Pattern Recognition · Computer Science 2019-06-18 Chen-Lin Zhang , Xin-Xin Liu , Jianxin Wu

Learning multimodal video understanding typically relies on datasets comprising video clips paired with manually annotated captions. However, this becomes even more challenging when dealing with long-form videos, lasting from minutes to…

Computer Vision and Pattern Recognition · Computer Science 2025-03-12 Soumya Shamarao Jahagirdar , Jayasree Saha , C V Jawahar

Multimodal Large Language Models (MLLMs) have demonstrated remarkable video reasoning capabilities across diverse tasks. However, their ability to understand human intent at a fine-grained level in egocentric videos remains largely…

Computer Vision and Pattern Recognition · Computer Science 2026-03-13 Ye Pan , Chi Kit Wong , Yuanhuiyi Lyu , Hanqian Li , Jiahao Huo , Jiacheng Chen , Lutao Jiang , Xu Zheng , Xuming Hu

Foundation models (FMs) are large neural networks trained on broad datasets, excelling in downstream tasks with minimal fine-tuning. Human activity recognition in video has advanced with FMs, driven by competition among different…

Computer Vision and Pattern Recognition · Computer Science 2024-07-23 Thinesh Thiyakesan Ponbagavathi , Kunyu Peng , Alina Roitberg

We propose a new dataset and a novel approach to learning hand-object interaction priors for hand and articulated object pose estimation. We first collect a dataset using visual teleoperation, where the human operator can directly play…

Computer Vision and Pattern Recognition · Computer Science 2024-07-30 Zehao Zhu , Jiashun Wang , Yuzhe Qin , Deqing Sun , Varun Jampani , Xiaolong Wang

This paper introduces a new video-and-language dataset with human actions for multimodal logical inference, which focuses on intentional and aspectual expressions that describe dynamic human actions. The dataset consists of 200 videos,…

Computer Vision and Pattern Recognition · Computer Science 2021-06-29 Riko Suzuki , Hitomi Yanaka , Koji Mineshima , Daisuke Bekki

Fingerspelling in sign language has been the means of communicating technical terms and proper nouns when they do not have dedicated sign language gestures. Automatic recognition of fingerspelling can help resolve communication barriers…

Computer Vision and Pattern Recognition · Computer Science 2021-08-24 Kamala Gajurel , Cuncong Zhong , Guanghui Wang

Tactile and visual perception are both crucial for humans to perform fine-grained interactions with their environment. Developing similar multi-modal sensing capabilities for robots can significantly enhance and expand their manipulation…

Robotics · Computer Science 2025-01-08 Binghao Huang , Yixuan Wang , Xinyi Yang , Yiyue Luo , Yunzhu Li

Few-shot action recognition aims to address the high cost and impracticality of manually labeling complex and variable video data in action recognition. It requires accurately classifying human actions in videos using only a few labeled…

Computer Vision and Pattern Recognition · Computer Science 2025-03-31 Yuyang Wanyan , Xiaoshan Yang , Weiming Dong , Changsheng Xu

We propose DeepMultiCap, a novel method for multi-person performance capture using sparse multi-view cameras. Our method can capture time varying surface details without the need of using pre-scanned template models. To tackle with the…

Computer Vision and Pattern Recognition · Computer Science 2021-08-31 Yang Zheng , Ruizhi Shao , Yuxiang Zhang , Tao Yu , Zerong Zheng , Qionghai Dai , Yebin Liu

Action recognition is so far mainly focusing on the problem of classification of hand selected preclipped actions and reaching impressive results in this field. But with the performance even ceiling on current datasets, it also appears that…

Computer Vision and Pattern Recognition · Computer Science 2019-06-05 Hilde Kuehne , Ahsan Iqbal , Alexander Richard , Juergen Gall

While fine-grained object recognition is an important problem in computer vision, current models are unlikely to accurately classify objects in the wild. These fully supervised models need additional annotated images to classify objects in…

Computer Vision and Pattern Recognition · Computer Science 2017-09-11 Timnit Gebru , Judy Hoffman , Li Fei-Fei

We address the problem of accurate capture of interactive behaviors between two people in daily scenarios. Most previous works either only consider one person or solely focus on conversational gestures of two people, assuming the body…

Computer Vision and Pattern Recognition · Computer Science 2025-09-09 Leo Ho , Yinghao Huang , Dafei Qin , Mingyi Shi , Wangpok Tse , Wei Liu , Junichi Yamagishi , Taku Komura

The improved competence of generative models can help building multi-modal virtual assistants that leverage modalities beyond language. By observing humans performing multi-step tasks, one can build assistants that have situational…

Computer Vision and Pattern Recognition · Computer Science 2025-01-22 Pha Nguyen , Sailik Sengupta , Girik Malik , Arshit Gupta , Bonan Min

This study delves into the realm of multi-modality (i.e., video and motion modalities) human behavior understanding by leveraging the powerful capabilities of Large Language Models (LLMs). Diverging from recent LLMs designed for video-only…

Computer Vision and Pattern Recognition · Computer Science 2024-05-31 Ling-Hao Chen , Shunlin Lu , Ailing Zeng , Hao Zhang , Benyou Wang , Ruimao Zhang , Lei Zhang