中文
相关论文

相关论文: ActionArt: Advancing Multimodal Large Models for F…

200 篇论文

Annotating videos is cumbersome, expensive and not scalable. Yet, many strong video models still rely on manually annotated data. With the recent introduction of the HowTo100M dataset, narrated videos now offer the possibility of learning…

计算机视觉与模式识别 · 计算机科学 2020-08-25 Antoine Miech , Jean-Baptiste Alayrac , Lucas Smaira , Ivan Laptev , Josef Sivic , Andrew Zisserman

Large language models (LLMs) have demonstrated exceptional capabilities in text understanding, which has paved the way for their expansion into video LLMs (Vid-LLMs) to analyze video data. However, current Vid-LLMs struggle to…

计算机视觉与模式识别 · 计算机科学 2026-02-04 Ming Nie , Dan Ding , Chunwei Wang , Yuanfan Guo , Jianhua Han , Hang Xu , Li Zhang

Classifying the behavior of humans or animals from videos is important in biomedical fields for understanding brain function and response to stimuli. Action recognition, classifying activities performed by one or more subjects in a trimmed…

计算机视觉与模式识别 · 计算机科学 2023-01-18 Michael Perez , Corey Toler-Franklin

State-of-the-art multimodal web agents, powered by Multimodal Large Language Models (MLLMs), can autonomously execute many web tasks by processing user instructions and interacting with graphical user interfaces (GUIs). Current strategies…

人工智能 · 计算机科学 2024-11-21 Gaurav Verma , Rachneet Kaur , Nishan Srishankar , Zhen Zeng , Tucker Balch , Manuela Veloso

When people observe and interact with physical spaces, they are able to associate functionality to regions in the environment. Our goal is to automate dense functional understanding of large spaces by leveraging sparse activity…

计算机视觉与模式识别 · 计算机科学 2016-05-06 Nicholas Rhinehart , Kris M. Kitani

Inspired by recent advances in neural machine translation, that jointly align and translate using encoder-decoder networks equipped with attention, we propose an attentionbased LSTM model for human activity recognition. Our model jointly…

计算机视觉与模式识别 · 计算机科学 2017-09-01 Atousa Torabi , Leonid Sigal

It is natural to represent objects in terms of their parts. This has the potential to improve the performance of algorithms for object recognition and segmentation but can also help for downstream tasks like activity recognition. Research…

计算机视觉与模式识别 · 计算机科学 2022-12-20 Ju He , Shuo Yang , Shaokang Yang , Adam Kortylewski , Xiaoding Yuan , Jie-Neng Chen , Shuai Liu , Cheng Yang , Qihang Yu , Alan Yuille

Multimodal Large Language Models (MLLMs) have shown remarkable capabilities in video content understanding but still struggle with fine-grained motion comprehension. To comprehensively assess the motion understanding ability of existing…

计算机视觉与模式识别 · 计算机科学 2025-03-20 Chongjun Tu , Lin Zhang , Pengtao Chen , Peng Ye , Xianfang Zeng , Wei Cheng , Gang Yu , Tao Chen

In egocentric action recognition a single population model is typically trained and subsequently embodied on a head-mounted device, such as an augmented reality headset. While this model remains static for new users and environments, we…

计算机视觉与模式识别 · 计算机科学 2023-07-13 Matthias De Lange , Hamid Eghbalzadeh , Reuben Tan , Michael Iuzzolino , Franziska Meier , Karl Ridgeway

We present a unified perspective on tackling various human-centric video tasks by learning human motion representations from large-scale and heterogeneous data resources. Specifically, we propose a pretraining stage in which a motion…

计算机视觉与模式识别 · 计算机科学 2023-08-15 Wentao Zhu , Xiaoxuan Ma , Zhaoyang Liu , Libin Liu , Wayne Wu , Yizhou Wang

We explore the human motion knowledge of Large Language Models (LLMs) through 3D avatar control. Given a motion instruction, we prompt LLMs to first generate a high-level movement plan with consecutive steps (High-level Planning), then…

计算机视觉与模式识别 · 计算机科学 2025-09-23 Kunhang Li , Jason Naradowsky , Yansong Feng , Yusuke Miyao

Human action analysis and understanding in videos is an important and challenging task. Although substantial progress has been made in past years, the explainability of existing methods is still limited. In this work, we propose a novel…

计算机视觉与模式识别 · 计算机科学 2019-08-29 Tao Zhuo , Zhiyong Cheng , Peng Zhang , Yongkang Wong , Mohan Kankanhalli

Vision-Language Models (VLMs) have achieved impressive progress in perceiving and describing visual environments. However, their ability to proactively reason and act based solely on visual inputs, without explicit textual prompts, remains…

计算机视觉与模式识别 · 计算机科学 2026-01-06 Daoan Zhang , Pai Liu , Xiaofei Zhou , Yuan Ge , Guangchen Lan , Jing Bi , Christopher Brinton , Ehsan Hoque , Jiebo Luo

Recognizing and categorizing human actions is an important task with applications in various fields such as human-robot interaction, video analysis, surveillance, video retrieval, health care system and entertainment industry. This thesis…

计算机视觉与模式识别 · 计算机科学 2021-05-03 Zahra Gharaee

Spatio-temporal action detection is an important and challenging problem in video understanding. However, the application of the existing large-scale spatio-temporal action datasets in specific fields is limited, and there is currently no…

计算机视觉与模式识别 · 计算机科学 2022-04-27 Fan Yang

Action recognition from videos, i.e., classifying a video into one of the pre-defined action types, has been a popular topic in the communities of artificial intelligence, multimedia, and signal processing. However, existing methods usually…

计算机视觉与模式识别 · 计算机科学 2022-09-05 Xiaodong Chen , Xinchen Liu , Wu Liu , Kun Liu , Dong Wu , Yongdong Zhang , Tao Mei

Despite substantial progress in video understanding, most existing datasets are limited to Earth's gravitational conditions. However, microgravity alters human motion, interactions, and visual semantics, revealing a critical gap for…

计算机视觉与模式识别 · 计算机科学 2026-03-24 Di Wen , Lei Qi , Kunyu Peng , Kailun Yang , Fei Teng , Ao Luo , Jia Fu , Yufan Chen , Ruiping Liu , Yitian Shi , M. Saquib Sarfraz , Rainer Stiefelhagen

Fine-grained video classification requires understanding complex spatio-temporal and semantic cues that often exceed the capacity of a single modality. In this paper, we propose a multimodal framework that fuses video, image, and text…

计算机视觉与模式识别 · 计算机科学 2025-07-08 Namho Kim , Junhwa Kim

Human-centric visual understanding is an important desideratum for effective human-robot interaction. In order to navigate crowded public places, social robots must be able to interpret the activity of the surrounding humans. This paper…

计算机视觉与模式识别 · 计算机科学 2023-07-28 Shengnan Hu , Ce Zheng , Zixiang Zhou , Chen Chen , Gita Sukthankar

Progress in embodied intelligence increasingly depends on scalable data infrastructure. While vision and language have scaled with internet corpora, learning physical interaction remains constrained by the lack of large, diverse, and richly…

计算机视觉与模式识别 · 计算机科学 2026-05-11 Yufan Deng , Daquan Zhou