English
Related papers

Related papers: FineParser: A Fine-grained Spatio-temporal Action …

200 papers

The development of multimodal large language models (MLLMs) enables the evaluation of image quality through natural language descriptions. This advancement allows for more detailed assessments. However, these MLLM-based IQA methods…

Computer Vision and Pattern Recognition · Computer Science 2026-03-03 Zheng Chen , Xun Zhang , Wenbo Li , Renjing Pei , Fenglong Song , Xiongkuo Min , Xiaohong Liu , Xin Yuan , Yong Guo , Yulun Zhang

What is the right way to reason about human activities? What directions forward are most promising? In this work, we analyze the current state of human activity understanding in videos. The goal of this paper is to examine datasets,…

Computer Vision and Pattern Recognition · Computer Science 2017-08-10 Gunnar A. Sigurdsson , Olga Russakovsky , Abhinav Gupta

Weakly-Supervised Video Anomaly Detection aims to identify anomalous events using only video-level labels, balancing annotation efficiency with practical applicability. However, existing methods often oversimplify the anomaly space by…

Computer Vision and Pattern Recognition · Computer Science 2025-12-30 Junhee Lee , ChaeBeen Bang , MyoungChul Kim , MyeongAh Cho

Actions are more than just movements and trajectories: we cook to eat and we hold a cup to drink from it. A thorough understanding of videos requires going beyond appearance modeling and necessitates reasoning about the sequence of…

Computer Vision and Pattern Recognition · Computer Science 2017-07-25 Gunnar A. Sigurdsson , Santosh Divvala , Ali Farhadi , Abhinav Gupta

In this paper, a novel signature of human action recognition, namely the curvature of a video sequence, is introduced. In this way, the distribution of sequential data is modeled, which enables few-shot learning. Instead of depending on…

Computer Vision and Pattern Recognition · Computer Science 2019-06-18 He Chen , Gregory S. Chirikjian

Fine-grained video classification requires understanding complex spatio-temporal and semantic cues that often exceed the capacity of a single modality. In this paper, we propose a multimodal framework that fuses video, image, and text…

Computer Vision and Pattern Recognition · Computer Science 2025-07-08 Namho Kim , Junhwa Kim

With the rapid progress of deepfake techniques in recent years, facial video forgery can generate highly deceptive video contents and bring severe security threats. And detection of such forgery videos is much more urgent and challenging.…

Computer Vision and Pattern Recognition · Computer Science 2021-06-25 Wei Lu , Lingyi Liu , Junwei Luo , Xianfeng Zhao , Yicong Zhou , Jiwu Huang

Composed Video Retrieval (CoVR) retrieves a target video given a query video and a modification text describing the intended change. Existing CoVR benchmarks emphasize appearance shifts or coarse event changes and therefore do not test the…

Computer Vision and Pattern Recognition · Computer Science 2025-11-21 Animesh Gupta , Jay Parmar , Ishan Rajendrakumar Dave , Mubarak Shah

Deepfake videos are causing growing concerns among communities due to their ever-increasing realism. Naturally, automated detection of forged Deepfake videos is attracting a proportional amount of interest of researchers. Current methods…

Computer Vision and Pattern Recognition · Computer Science 2025-02-13 Yunzhuo Chen , Naveed Akhtar , Nur Al Hasan Haldar , Ajmal Mian

Recent advancements in personalized Text-to-Video (T2V) generation have made significant strides in synthesizing character-specific content. However, these methods face a critical limitation: the inability to perform fine-grained control…

Computer Vision and Pattern Recognition · Computer Science 2026-01-06 Haopeng Fang , Di Qiu , Binjie Mao , He Tang

Advancements in attention mechanisms have led to significant performance improvements in a variety of areas in machine learning due to its ability to enable the dynamic modeling of temporal sequences. A particular area in computer vision…

Computer Vision and Pattern Recognition · Computer Science 2021-12-14 Brennan Gebotys , Alexander Wong , David A. Clausi

Vision-based human activity recognition has emerged as one of the essential research areas in video analytics domain. Over the last decade, numerous advanced deep learning algorithms have been introduced to recognize complex human actions…

Computer Vision and Pattern Recognition · Computer Science 2022-08-11 Hayat Ullah , Arslan Munir

We present a dual-pathway approach for recognizing fine-grained interactions from videos. We build on the success of prior dual-stream approaches, but make a distinction between the static and dynamic representations of objects and their…

Computer Vision and Pattern Recognition · Computer Science 2021-04-02 Tae Soo Kim , Jonathan Jones , Gregory D. Hager

Human Action Recognition (HAR) aims to understand human behavior and assign a label to each action. It has a wide range of applications, and therefore has been attracting increasing attention in the field of computer vision. Human actions…

Computer Vision and Pattern Recognition · Computer Science 2022-06-22 Zehua Sun , Qiuhong Ke , Hossein Rahmani , Mohammed Bennamoun , Gang Wang , Jun Liu

Video quality assessment (VQA) has attracted growing attention in recent years. While the great expense of annotating large-scale VQA datasets has become the main obstacle for current deep-learning methods. To surmount the constraint of…

Computer Vision and Pattern Recognition · Computer Science 2023-08-03 Hongbo Liu , Mingda Wu , Kun Yuan , Ming Sun , Yansong Tang , Chuanchuan Zheng , Xing Wen , Xiu Li

Action quality assessment (AQA) aims to automatically quantify the execution quality of human actions in videos and is valuable for applications such as competitive sports judging. In multimodal AQA, quality evidence from different…

Computer Vision and Pattern Recognition · Computer Science 2026-05-12 Qiqi Li , Pengfei Wang , Nenggan Zheng

Text-to-video generation has shown promising results. However, by taking only natural languages as input, users often face difficulties in providing detailed information to precisely control the model's output. In this work, we propose…

Computer Vision and Pattern Recognition · Computer Science 2023-12-06 Hsin-Ping Huang , Yu-Chuan Su , Deqing Sun , Lu Jiang , Xuhui Jia , Yukun Zhu , Ming-Hsuan Yang

We propose a novel method for temporally pooling frames in a video for the task of human action recognition. The method is motivated by the observation that there are only a small number of frames which, together, contain sufficient…

Computer Vision and Pattern Recognition · Computer Science 2017-06-27 Amlan Kar , Nishant Rai , Karan Sikka , Gaurav Sharma

Video Question Answering (VideoQA) based on Large Language Models (LLMs) has shown potential in general video understanding but faces significant challenges when applied to the inherently complex domain of sports videos. In this work, we…

Computer Vision and Pattern Recognition · Computer Science 2025-09-16 Haodong Chen , Haojian Huang , XinXiang Yin , Dian Shao

Affect is often expressed via non-verbal body language such as actions/gestures, which are vital indicators for human behaviors. Recent studies on recognition of fine-grained actions/gestures in monocular images have mainly focused on…

Computer Vision and Pattern Recognition · Computer Science 2021-01-19 Ardhendu Behera , Zachary Wharton , Morteza Ghahremani , Swagat Kumar , Nik Bessis
‹ Prev 1 8 9 10 Next ›