中文
相关论文

相关论文: Skeleton-Guided Spatial-Temporal Feature Learning …

200 篇论文

It is challenging to annotate large-scale datasets for supervised video shadow detection methods. Using a model trained on labeled images to the video frames directly may lead to high generalization error and temporal inconsistent results.…

计算机视觉与模式识别 · 计算机科学 2022-06-20 Xiao Lu , Yihong Cao , Sheng Liu , Chengjiang Long , Zipei Chen , Xuanyu Zhou , Yimin Yang , Chunxia Xiao

Session-Based Recommenders (SBRs) aim to predict users' next preferences regard to their previous interactions in sessions while there is no historical information about them. Modern SBRs utilize deep neural networks to map users' current…

信息检索 · 计算机科学 2023-12-18 Reza Yeganegi , Saman Haratizadeh

Human action recognition is crucial in computer vision systems. However, in real-world scenarios, human actions often fall outside the distribution of training data, requiring a model to both recognize in-distribution (ID) actions and…

计算机视觉与模式识别 · 计算机科学 2024-12-20 Jing Xu , Anqi Zhu , Jingyu Lin , Qiuhong Ke , Cunjian Chen

Contemporary Video Instance Segmentation (VIS) methods typically adhere to a pre-train then fine-tune regime, where a segmentation model trained on images is fine-tuned on videos. However, the lack of temporal knowledge in the pre-trained…

计算机视觉与模式识别 · 计算机科学 2025-03-25 Qing Zhong , Peng-Tao Jiang , Wen Wang , Guodong Ding , Lin Wu , Kaiqi Huang

Video super-resolution (VSR) can achieve better performance compared to single image super-resolution by additionally leveraging temporal information. In particular, the recurrent-based VSR model exploits long-range temporal information…

计算机视觉与模式识别 · 计算机科学 2025-05-06 Xingyu Zhou , Wei Long , Jingbo Lu , Shiyin Jiang , Weiyi You , Haifeng Wu , Shuhang Gu

Detecting human-object interactions (HOI) is an important step toward a comprehensive visual understanding of machines. While detecting non-temporal HOIs (e.g., sitting on a chair) from static images is feasible, it is unlikely even for…

计算机视觉与模式识别 · 计算机科学 2021-06-25 Meng-Jiun Chiou , Chun-Yu Liao , Li-Wei Wang , Roger Zimmermann , Jiashi Feng

Weakly supervised instance segmentation reduces the cost of annotations required to train models. However, existing approaches which rely only on image-level class labels predominantly suffer from errors due to (a) partial segmentation of…

计算机视觉与模式识别 · 计算机科学 2021-03-25 Qing Liu , Vignesh Ramanathan , Dhruv Mahajan , Alan Yuille , Zhenheng Yang

With rich temporal-spatial information, video-based person re-identification methods have shown broad prospects. Although tracklets can be easily obtained with ready-made tracking models, annotating identities is still expensive and…

计算机视觉与模式识别 · 计算机科学 2026-01-21 Nanxing Meng , Qizao Wang , Bin Li , Xiangyang Xue

Vehicle re-identification (ReID) in a large-scale camera network is important in public safety, traffic control, and security. However, due to the appearance ambiguities of vehicle, the previous appearance-based ReID methods often fail to…

计算机视觉与模式识别 · 计算机科学 2023-09-06 Hye-Geun Kim , YouKyoung Na , Hae-Won Joe , Yong-Hyuk Moon , Yeong-Jun Cho

Video Question Answering (VideoQA) task serves as a critical playground for evaluating whether foundation models can effectively perceive, understand, and reason about dynamic real-world scenarios. However, existing Multimodal Large…

计算机视觉与模式识别 · 计算机科学 2025-12-12 Sunqi Fan , Jiashuo Cui , Meng-Hao Guo , Shuojin Yang

There exist a wide range of intra class variations of the same actions and inter class similarity among the actions, at the same time, which makes the action recognition in videos very challenging. In this paper, we present a novel…

计算机视觉与模式识别 · 计算机科学 2019-12-03 Chhavi Dhiman , Dinesh Kumar Vishwakarma , Paras Aggarwal

Skeleton-based temporal action segmentation is a fundamental yet challenging task, playing a crucial role in enabling intelligent systems to perceive and respond to human activities. While fully-supervised methods achieve satisfactory…

计算机视觉与模式识别 · 计算机科学 2026-03-09 Hongsong Wang , Yiqin Shen , Pengbo Yan , Jie Gui

Visible-infrared person re-identification (VI-ReID) is a task of matching the same individuals across the visible and infrared modalities. Its main challenge lies in the modality gap caused by cameras operating on different spectra.…

计算机视觉与模式识别 · 计算机科学 2022-08-23 Qiong Wu , Jiaer Xia , Pingyang Dai , Yiyi Zhou , Yongjian Wu , Rongrong Ji

Recent progress in Multimodal Large Language Models (MLLMs) has demonstrated strong semantic understanding capabilities, but struggles to perform precise spatio-temporal understanding. Existing spatio-temporal methods primarily focus on the…

Video-based gaze estimation methods aim to capture the inherently temporal dynamics of human eye gaze from multiple image frames. However, since models must capture both spatial and temporal relationships, performance is limited by the…

计算机视觉与模式识别 · 计算机科学 2025-12-22 Alexandre Personnic , Mihai Bâce

Scaling multimodal large language models (MLLMs) to long videos is constrained by limited context windows. While retrieval-augmented generation (RAG) is a promising remedy by organizing query-relevant visual evidence into a compact context,…

计算机视觉与模式识别 · 计算机科学 2026-04-17 Honghao Fu , Miao Xu , Yiwei Wang , Dailing Zhang , Jun Liu , Yujun Cai

Video Instance Segmentation (VIS) is a task that simultaneously requires classification, segmentation, and instance association in a video. Recent VIS approaches rely on sophisticated pipelines to achieve this goal, including RoI-related…

计算机视觉与模式识别 · 计算机科学 2022-08-23 Zhengkai Jiang , Zhangxuan Gu , Jinlong Peng , Hang Zhou , Liang Liu , Yabiao Wang , Ying Tai , Chengjie Wang , Liqing Zhang

Neural Representations for Videos (NeRV) has emerged as a promising implicit neural representation (INR) approach for video analysis, which represents videos as neural networks with frame indexes as inputs. However, NeRV-based methods are…

计算机视觉与模式识别 · 计算机科学 2025-01-20 Jialong Guo , Ke liu , Jiangchao Yao , Zhihua Wang , Jiajun Bu , Haishuai Wang

Visual Autoregressive modeling (VAR) has emerged as a highly efficient alternative to diffusion-based frameworks, achieving comparable synthesis quality. However, as this paradigm extends to Spacetime Autoregressive modeling (STAR) for…

计算机视觉与模式识别 · 计算机科学 2026-03-10 Sungwoong Yune , Suheon Jeong , Joo-Young Kim

Recently, visual-language learning (VLL) has shown great potential in enhancing visual-based person re-identification (ReID). Existing VLL-based ReID methods typically focus on image-text feature alignment at the whole-body level, while…

计算机视觉与模式识别 · 计算机科学 2025-03-24 Yin Lin , Yehansen Chen , Baocai Yin , Jinshui Hu , Bing Yin , Cong Liu , Zengfu Wang