中文
相关论文

相关论文: Multiresolution Match Kernels for Gesture Video Cl…

200 篇论文

Several recent works have directly extended the image masked autoencoder (MAE) with random masking into video domain, achieving promising results. However, unlike images, both spatial and temporal information are important for video…

计算机视觉与模式识别 · 计算机科学 2023-08-25 David Fan , Jue Wang , Shuai Liao , Yi Zhu , Vimal Bhat , Hector Santos-Villalobos , Rohith MV , Xinyu Li

In this paper, we present the Bag-of-Attributes (BoA) model for video representation aiming at video event retrieval. The BoA model is based on a semantic feature space for representing videos, resulting in high-level video feature vectors.…

信息检索 · 计算机科学 2020-12-29 Leonardo A. Duarte , Otávio A. B. Penatti , Jurandy Almeida

With the growing amount of inappropriate content on the Internet, such as pornography, arises the need to detect and filter such material. The reason for this is given by the fact that such content is often prohibited in certain…

计算机视觉与模式识别 · 计算机科学 2016-11-14 Carlos Caetano , Sandra Avila , William Robson Schwartz , Silvio Jamil F. Guimarães , Arnaldo de A. Araújo

Motion estimation is one of the important procedures in the all video encoders. Most of the complexity of the video coder depends on the complexity of the motion estimation step. The original motion estimation algorithm has a remarkable…

图像与视频处理 · 电气工程与系统科学 2018-03-14 Amin Banitalebi , Said Nader-Esfahani , Alireza Nasiri Avanaki

We present an on-device real-time hand gesture recognition (HGR) system, which detects a set of predefined static gestures from a single RGB camera. The system consists of two parts: a hand skeleton tracker and a gesture classifier. We use…

In this paper, we propose a method for image-set classification based on convex cone models. Image set classification aims to classify a set of images, which were usually obtained from video frames or multi-view cameras, into a target…

计算机视觉与模式识别 · 计算机科学 2019-03-18 Naoya Sogi , Rui Zhu , Jing-Hao Xue , Kazuhiro Fukui

Object-level spatial-temporal understanding is essential for video question answering, yet existing multimodal large language models (MLLMs) encode frames holistically and lack explicit mechanisms for fine-grained object grounding. Recent…

计算机视觉与模式识别 · 计算机科学 2026-04-14 Zekun Qian , Ruize Han , Wei Feng

Video generation models have developed rapidly in recent years, where generating natural human motion plays a pivotal role. However, accurately evaluating the quality of generated human motion video remains a significant challenge. Existing…

计算机视觉与模式识别 · 计算机科学 2026-04-29 Bingzi Zhang , Kaisi Guan , Ruihua Song

In this paper, a real-time signal processing frame-work based on a 60 GHz frequency-modulated continuous wave (FMCW) radar system to recognize gestures is proposed. In order to improve the robustness of the radar-based gesture recognition…

信号处理 · 电气工程与系统科学 2020-05-21 Yuliang Sun , Tai Fei , Xibo Li , Alexander Warnecke , Ernst Warsitz , Nils Pohl

Multimodal synthetic data generation is crucial in domains such as autonomous driving, robotics, augmented/virtual reality, and retail. We propose a novel approach, GenMM, for jointly editing RGB videos and LiDAR scans by inserting…

计算机视觉与模式识别 · 计算机科学 2024-06-18 Bharat Singh , Viveka Kulharia , Luyu Yang , Avinash Ravichandran , Ambrish Tyagi , Ashish Shrivastava

Continuous mid-air hand gesture recognition based on captured hand pose streams is fundamental for human-computer interaction, particularly in AR / VR. However, many of the methods proposed to recognize heterogeneous hand gestures are…

计算机视觉与模式识别 · 计算机科学 2023-04-13 Federico Cunico , Federico Girella , Andrea Avogaro , Marco Emporio , Andrea Giachetti , Marco Cristani

We propose a novel scene-segmentation-based exposure compensation method for multi-exposure image fusion (MEF) based tone mapping. The aim of MEF-based tone mapping is to display high dynamic range (HDR) images on devices with limited…

计算机视觉与模式识别 · 计算机科学 2024-10-29 Yuma Kinoshita , Hitoshi Kiya

Recent deep learning approaches have achieved impressive performance on visual sound separation tasks. However, these approaches are mostly built on appearance and optical flow like motion feature representations, which exhibit limited…

计算机视觉与模式识别 · 计算机科学 2020-04-21 Chuang Gan , Deng Huang , Hang Zhao , Joshua B. Tenenbaum , Antonio Torralba

The recognition of behaviors in videos usually requires a combinatorial analysis of the spatial information about objects and their dynamic action information in the temporal dimension. Specifically, behavior recognition may even rely more…

计算机视觉与模式识别 · 计算机科学 2022-03-08 Lizong Zhang , Yiming Wang , Bei Hui , Xiujian Zhang , Sijuan Liu , Shuxin Feng

Many of the state-of-the-art algorithms for gesture recognition are based on Conditional Random Fields (CRFs). Successful approaches, such as the Latent-Dynamic CRFs, extend the CRF by incorporating latent variables, whose values are mapped…

计算机视觉与模式识别 · 计算机科学 2017-04-04 Manoel Horta Ribeiro , Bruno Teixeira , Antônio Otávio Fernandes , Wagner Meira , Erickson R. Nascimento

This paper introduces a state-of-the-art video representation and applies it to efficient action recognition and detection. We first propose to improve the popular dense trajectory features by explicit camera motion estimation. More…

计算机视觉与模式识别 · 计算机科学 2015-04-22 Heng Wang , Dan Oneata , Jakob Verbeek , Cordelia Schmid

HMMs are widely used in action and gesture recognition due to their implementation simplicity, low computational requirement, scalability and high parallelism. They have worth performance even with a limited training set. All these…

计算机视觉与模式识别 · 计算机科学 2017-03-09 Guido Borghi , Roberto Vezzani , Rita Cucchiara

Recognizing human actions in untrimmed videos is an important challenging task. An effective 3D motion representation and a powerful learning model are two key factors influencing recognition performance. In this paper we introduce a new…

计算机视觉与模式识别 · 计算机科学 2018-12-31 Huy-Hieu Pham , Louahdi Khoudour , Alain Crouzil , Pablo Zegers , Sergio A. Velastin

Human actions in video sequences are characterized by the complex interplay between spatial features and their temporal dynamics. In this paper, we propose novel tensor representations for compactly capturing such higher-order relationships…

计算机视觉与模式识别 · 计算机科学 2021-08-31 Piotr Koniusz , Lei Wang , Anoop Cherian

Hand gestures form an intuitive means of interaction in Mixed Reality (MR) applications. However, accurate gesture recognition can be achieved only through state-of-the-art deep learning models or with the use of expensive sensors. Despite…

计算机视觉与模式识别 · 计算机科学 2019-04-23 Varun Jain , Gaurav Garg , Ramakrishna Perla , Ramya Hebbalaguppe