English
Related papers

Related papers: D3D: Distilled 3D Networks for Video Action Recogn…

200 papers

This paper presents a novel spatiotemporal transformer network that introduces several original components to detect actions in untrimmed videos. First, the multi-feature selective semantic attention model calculates the correlations…

Computer Vision and Pattern Recognition · Computer Science 2024-05-15 Matthew Korban , Peter Youngs , Scott T. Acton

We investigate video classification via a two-stream convolutional neural network (CNN) design that directly ingests information extracted from compressed video bitstreams. Our approach begins with the observation that all modern video…

Computer Vision and Pattern Recognition · Computer Science 2017-12-21 Aaron Chadha , Alhabib Abbas , Yiannis Andreopoulos

The landscape of video recognition has evolved significantly, shifting from traditional Convolutional Neural Networks (CNNs) to Transformer-based architectures for improved accuracy. While 3D CNNs have been effective at capturing…

Computer Vision and Pattern Recognition · Computer Science 2025-07-21 Hayat Ullah , Muhammad Ali Shafique , Abbas Khan , Arslan Munir

3D Convolutional Neural Network (3D CNN) captures spatial and temporal information on 3D data such as video sequences. However, due to the convolution and pooling mechanism, the information loss seems unavoidable. To improve the visual…

Computer Vision and Pattern Recognition · Computer Science 2022-08-17 Novanto Yudistira , Muthu Subash Kavitha , Takio Kurita

The spatio-temporal information among video sequences is significant for video super-resolution (SR). However, the spatio-temporal information cannot be fully used by existing video SR methods since spatial feature extraction and temporal…

Computer Vision and Pattern Recognition · Computer Science 2021-11-29 Xinyi Ying , Longguang Wang , Yingqian Wang , Weidong Sheng , Wei An , Yulan Guo

We present a self-supervised approach using spatio-temporal signals between video frames for action recognition. A two-stream architecture is leveraged to tangle spatial and temporal representation learning. Our task is formulated as both a…

Computer Vision and Pattern Recognition · Computer Science 2018-06-20 Ahmed Taha , Moustafa Meshry , Xitong Yang , Yi-Ting Chen , Larry Davis

This paper aims to accelerate video stream processing, such as object detection and semantic segmentation, by leveraging the temporal redundancies that exist between video frames. Instead of propagating and warping features using motion…

Computer Vision and Pattern Recognition · Computer Science 2022-03-21 Amirhossein Habibian , Haitam Ben Yahia , Davide Abati , Efstratios Gavves , Fatih Porikli

In this paper, several variants of two-stream architectures for temporal action proposal generation in long, untrimmed videos are presented. Inspired by the recent advances in the field of human action recognition utilizing 3D convolutions…

Computer Vision and Pattern Recognition · Computer Science 2019-03-15 Patrick Schlosser , David Münch , Michael Arens

The performance of video saliency estimation techniques has achieved significant advances along with the rapid development of Convolutional Neural Networks (CNNs). However, devices like cameras and drones may have limited computational…

Computer Vision and Pattern Recognition · Computer Science 2020-01-08 Jia Li , Kui Fu , Shengwei Zhao , Shiming Ge

Video classification is productive in many practical applications, and the recent deep learning has greatly improved its accuracy. However, existing works often model video frames indiscriminately, but from the view of motion, video frames…

Computer Vision and Pattern Recognition · Computer Science 2017-03-28 Yunzhen Zhao , Yuxin Peng

Understanding actions and gestures in video streams requires temporal reasoning of the spatial content from different time instants, i.e., spatiotemporal (ST) modeling. In this survey paper, we have made a comparative analysis of different…

Computer Vision and Pattern Recognition · Computer Science 2021-01-12 Okan Köpüklü , Fabian Herzog , Gerhard Rigoll

In this paper, we address the challenging problem of action recognition, using event-based cameras. To recognise most gestural actions, often higher temporal precision is required for sampling visual information. Actions are defined by…

Computer Vision and Pattern Recognition · Computer Science 2019-03-19 Rohan Ghosh , Anupam Gupta , Andrei Nakagawa , Alcimar Soares , Nitish Thakor

In this thesis, we focus on video action understanding problems from an online and real-time processing point of view. We start with the conversion of the traditional offline spatiotemporal action detection pipeline into an online…

Computer Vision and Pattern Recognition · Computer Science 2020-09-01 Gurkirt Singh

The existing action recognition methods are mainly based on clip-level classifiers such as two-stream CNNs or 3D CNNs, which are trained from the randomly selected clips and applied to densely sampled clips during testing. However, this…

Computer Vision and Pattern Recognition · Computer Science 2020-08-26 Yin-Dong Zheng , Zhaoyang Liu , Tong Lu , Limin Wang

Video Analytics Software as a Service (VA SaaS) has been rapidly growing in recent years. VA SaaS is typically accessed by users using a lightweight client. Because the transmission bandwidth between the client and cloud is usually limited…

Computer Vision and Pattern Recognition · Computer Science 2018-08-16 Zhaoyang Zhang , Zhanghui Kuang , Ping Luo , Litong Feng , Wei Zhang

In this work, we introduce a new video representation for action classification that aggregates local convolutional features across the entire spatio-temporal extent of the video. We do so by integrating state-of-the-art two-stream networks…

Computer Vision and Pattern Recognition · Computer Science 2017-04-11 Rohit Girdhar , Deva Ramanan , Abhinav Gupta , Josef Sivic , Bryan Russell

Video classification is highly important with wide applications, such as video search and intelligent surveillance. Video naturally consists of static and motion information, which can be represented by frame and optical flow. Recently,…

Computer Vision and Pattern Recognition · Computer Science 2017-11-10 Yuxin Peng , Yunzhen Zhao , Junchao Zhang

We focus on the word-level visual lipreading, which requires recognizing the word being spoken, given only the video but not the audio. State-of-the-art methods explore the use of end-to-end neural networks, including a shallow (up to three…

Computer Vision and Pattern Recognition · Computer Science 2019-07-22 Xinshuo Weng , Kris Kitani

This paper describes a network that captures multimodal correlations over arbitrary timestamps. The proposed scheme operates as a complementary, extended network over a multimodal convolutional neural network (CNN). Spatial and temporal…

Computer Vision and Pattern Recognition · Computer Science 2019-12-17 Novanto Yudistira , Takio Kurita

Recently, dataset distillation has paved the way towards efficient machine learning, especially for image datasets. However, the distillation for videos, characterized by an exclusive temporal dimension, remains an underexplored domain. In…

Computer Vision and Pattern Recognition · Computer Science 2024-04-16 Ziyu Wang , Yue Xu , Cewu Lu , Yong-Lu Li
‹ Prev 1 3 4 5 6 7 10 Next ›