English
Related papers

Related papers: Learning Mid-level Words on Riemannian Manifold fo…

200 papers

Representing videos by densely extracted local space-time features has recently become a popular approach for analysing actions. In this paper, we tackle the problem of categorising human actions by devising Bag of Words (BoW) models based…

Computer Vision and Pattern Recognition · Computer Science 2016-07-08 Masoud Faraki , Maziar Palhang , Conrad Sanderson

This paper addresses the problem of learning word image representations: given the cropped image of a word, we are interested in finding a descriptive, robust, and compact fixed-length representation. Machine learning techniques can then be…

Computer Vision and Pattern Recognition · Computer Science 2014-11-17 Albert Gordo

Realistic videos of human actions exhibit rich spatiotemporal structures at multiple levels of granularity: an action can always be decomposed into multiple finer-grained elements in both space and time. To capture this intuition, we…

Computer Vision and Pattern Recognition · Computer Science 2015-09-01 Tian Lan , Yuke Zhu , Amir Roshan Zamir , Silvio Savarese

Sparsity-based representations have recently led to notable results in various visual recognition tasks. In a separate line of research, Riemannian manifolds have been shown useful for dealing with features and models that do not lie in…

Machine Learning · Computer Science 2015-05-21 Mehrtash Harandi , Richard Hartley , Chunhua Shen , Brian Lovell , Conrad Sanderson

The importance of wild video based image set recognition is becoming monotonically increasing. However, the contents of these collected videos are often complicated, and how to efficiently perform set modeling and feature extraction is a…

Computer Vision and Pattern Recognition · Computer Science 2019-08-07 Rui Wang , XiaoJun Wu , Josef Kittler

We present a comparative evaluation of various techniques for action recognition while keeping as many variables as possible controlled. We employ two categories of Riemannian manifolds: symmetric positive definite matrices and linear…

Computer Vision and Pattern Recognition · Computer Science 2016-10-05 Johanna Carvajal , Arnold Wiliem , Chris McCool , Brian Lovell , Conrad Sanderson

Visual Recognition is one of the fundamental challenges in AI, where the goal is to understand the semantics of visual data. Employing mid-level representation, in particular, shifted the paradigm in visual recognition. The mid-level…

Computer Vision and Pattern Recognition · Computer Science 2015-12-24 Moin Nabi

This paper proposes a novel latent semantic learning method for extracting high-level features (i.e. latent semantics) from a large vocabulary of abundant mid-level features (i.e. visual keywords) with structured sparse representation,…

Multimedia · Computer Science 2015-03-19 Zhiwu Lu , Yuxin Peng

Riemannian manifolds have been widely employed for video representations in visual classification tasks including video-based face recognition. The success mainly derives from learning a discriminant Riemannian metric which encodes the…

Computer Vision and Pattern Recognition · Computer Science 2017-01-10 Zhiwu Huang , Ruiping Wang , Shiguang Shan , Luc Van Gool , Xilin Chen

Exploring open-vocabulary video action recognition is a promising venture, which aims to recognize previously unseen actions within any arbitrary set of categories. Existing methods typically adapt pretrained image-text models to the video…

Computer Vision and Pattern Recognition · Computer Science 2023-12-06 Chengyou Jia , Minnan Luo , Xiaojun Chang , Zhuohang Dang , Mingfei Han , Mengmeng Wang , Guang Dai , Sizhe Dang , Jingdong Wang

In this work\footnote {This work was supported in part by the National Science Foundation under grant IIS-1212948.}, we present a method to represent a video with a sequence of words, and learn the temporal sequencing of such words as the…

Computer Vision and Pattern Recognition · Computer Science 2019-06-18 Sangwoo Cho , Hassan Foroosh

Euclidean representations distort data with intrinsic non-Euclidean structure. While Riemannian representation learning offers a solution by embedding data onto matching manifolds, it typically relies on an encoder to estimate densities on…

Machine Learning · Computer Science 2026-05-05 Andreas Bjerregaard , Søren Hauberg , Anders Krogh

The purpose of mid-level visual element discovery is to find clusters of image patches that are both representative and discriminative. Here we study this problem from the prospective of pattern mining while relying on the recently…

Computer Vision and Pattern Recognition · Computer Science 2016-05-31 Yao Li , Lingqiao Liu , Chunhua Shen , Anton van den Hengel

Existing deep learning methods for action recognition in videos require a large number of labeled videos for training, which is labor-intensive and time-consuming. For the same action, the knowledge learned from different media types, e.g.,…

Computer Vision and Pattern Recognition · Computer Science 2020-02-19 Yang Liu , Zhaoyang Lu , Jing Li , Tao Yang , Chao Yao

A large amount of recent research has focused on tasks that combine language and vision, resulting in a proliferation of datasets and methods. One such task is action recognition, whose applications include image annotation, scene under-…

Computation and Language · Computer Science 2017-04-25 Spandana Gella , Frank Keller

Despite an exciting new wave of multimodal machine learning models, current approaches still struggle to interpret the complex contextual relationships between the different modalities present in videos. Going beyond existing methods that…

Computer Vision and Pattern Recognition · Computer Science 2023-09-20 Laura Hanu , Anita L. Verő , James Thewlis

In image set classification, a considerable advance has been made by modeling the original image sets by second order statistics or linear subspace, which typically lie on the Riemannian manifold. Specifically, they are Symmetric Positive…

Computer Vision and Pattern Recognition · Computer Science 2018-05-31 Rui Wang , Xiao-Jun Wu , Kai-Xuan Chen , Josef Kittler

Multimodal Language Analysis is a demanding area of research, since it is associated with two requirements: combining different modalities and capturing temporal information. During the last years, several works have been proposed in the…

Computation and Language · Computer Science 2022-01-10 Panagiotis Koromilas , Theodoros Giannakopoulos

In this paper, we examined the zero-shot activity recognition task with the usage of videos. We introduce an auto-encoder based model to construct a multimodal joint embedding space between the visual and textual manifolds. On the visual…

Computer Vision and Pattern Recognition · Computer Science 2020-02-07 Evin Pinar Ornek

Action recognition is so far mainly focusing on the problem of classification of hand selected preclipped actions and reaching impressive results in this field. But with the performance even ceiling on current datasets, it also appears that…

Computer Vision and Pattern Recognition · Computer Science 2019-06-05 Hilde Kuehne , Ahsan Iqbal , Alexander Richard , Juergen Gall
‹ Prev 1 2 3 10 Next ›