English
Related papers

Related papers: Bag of Visual Words and Fusion Methods for Action …

200 papers

Feature fusion is a commonly used strategy in image retrieval tasks, which aggregates the matching responses of multiple visual features. Feasible sets of features can be either descriptors (SIFT, HSV) for an entire image or the same…

Information Retrieval · Computer Science 2018-11-01 Zhongdao Wang , Liang Zheng , Shengjin Wang

The characteristics of feature selection, nonlinear combination and multi-task auxiliary learning mechanism of the human visual perception system play an important role in real-world scenarios, but the research of image fusion theory based…

Computer Vision and Pattern Recognition · Computer Science 2020-06-23 Aiqing Fang , Xinbo Zhao , Jiaqi Yang , Yanning Zhang

Recognizing human activities in videos is challenging due to the spatio-temporal complexity and context-dependence of human interactions. Prior studies often rely on single input modalities, such as RGB or skeletal data, limiting their…

Computer Vision and Pattern Recognition · Computer Science 2024-09-05 Tuyen Tran , Thao Minh Le , Hung Tran , Truyen Tran

In this paper we revisit feature fusion, an old-fashioned topic, in the new context of text-to-video retrieval. Different from previous research that considers feature fusion only at one end, let it be video or text, we aim for feature…

Multimedia · Computer Science 2022-07-28 Fan Hu , Aozhu Chen , Ziyue Wang , Fangming Zhou , Jianfeng Dong , Xirong Li

Continuous action recognition is more challenging than isolated recognition because classification and segmentation must be simultaneously carried out. We build on the well known dynamic time warping (DTW) framework and devise a novel…

Computer Vision and Pattern Recognition · Computer Science 2015-03-30 Kaustubh Kulkarni , Georgios Evangelidis , Jan Cech , Radu Horaud

The recent trend in action recognition is towards larger datasets, an increasing number of action classes and larger visual vocabularies. State-of-the-art human action classification in challenging video data is currently based on a…

Computer Vision and Pattern Recognition · Computer Science 2014-05-30 Michael Sapienza , Fabio Cuzzolin , Philip H. S. Torr

Video content is rich in semantics and has the ability to evoke various emotions in viewers. In recent years, with the rapid development of affective computing and the explosive growth of visual data, affective video content analysis (AVCA)…

Computer Vision and Pattern Recognition · Computer Science 2024-01-19 Junxiao Xue , Jie Wang , Xuecheng Wu , Qian Zhang

In this paper, we propose a new multi-layer structural approach for the task of object based image retrieval. In our work we tackle the problem of structural organization of local features. The structural features we propose are nested…

Multimedia · Computer Science 2014-05-15 Svebor Karaman , Jenny Benois-Pineau , Rémi Mégret , Aurélie Bugeau

A complex visual navigation task puts an agent in different situations which call for a diverse range of visual perception abilities. For example, to "go to the nearest chair", the agent might need to identify a chair in a living room using…

Computer Vision and Pattern Recognition · Computer Science 2021-08-05 Bokui Shen , Danfei Xu , Yuke Zhu , Leonidas J. Guibas , Li Fei-Fei , Silvio Savarese

Learning from visual data opens the potential to accrue a large range of manipulation behaviors by leveraging human demonstrations without specifying each of them mathematically, but rather through natural task specification. In this paper,…

Robotics · Computer Science 2021-11-16 Haoyu Xiong , Quanzhou Li , Yun-Chun Chen , Homanga Bharadhwaj , Samarth Sinha , Animesh Garg

Video Object Segmentation (VOS) is one of the most fundamental and challenging tasks in computer vision and has a wide range of applications. Most existing methods rely on spatiotemporal memory networks to extract frame-level features and…

Computer Vision and Pattern Recognition · Computer Science 2025-04-15 Mengjiao Wang , Junpei Zhang , Xu Liu , Yuting Yang , Mengru Ma

Continuous Bag of Words (CBOW) is a powerful text embedding method. Due to its strong capabilities to encode word content, CBOW embeddings perform well on a wide range of downstream tasks while being efficient to compute. However, CBOW is…

Computation and Language · Computer Science 2019-02-19 Florian Mai , Lukas Galke , Ansgar Scherp

We present an end-to-end method for object detection and trajectory prediction utilizing multi-view representations of LiDAR returns and camera images. In this work, we recognize the strengths and weaknesses of different view…

Computer Vision and Pattern Recognition · Computer Science 2021-10-20 Sudeep Fadadu , Shreyash Pandey , Darshan Hegde , Yi Shi , Fang-Chieh Chou , Nemanja Djuric , Carlos Vallespi-Gonzalez

Extensive research efforts have been dedicated to 3D model retrieval in recent decades. Recently, view-based methods have attracted much research attention due to the high discriminative property of multi-views for 3D object representation.…

Computer Vision and Pattern Recognition · Computer Science 2012-08-20 Qiong Liu

Continuous emotion recognition in terms of valence and arousal under in-the-wild (ITW) conditions remains a challenging problem due to large variations in appearance, head pose, illumination, occlusions, and subject-specific patterns of…

Computer Vision and Pattern Recognition · Computer Science 2026-03-16 Elena Ryumina , Maxim Markitantov , Alexandr Axyonov , Dmitry Ryumin , Mikhail Dolgushin , Denis Dresvyanskiy , Alexey Karpov

The emergence of depth imaging technologies like the Microsoft Kinect has renewed interest in computational methods for gesture classification based on videos. For several years now, researchers have used the Bag-of-Features (BoF) as a…

Computer Vision and Pattern Recognition · Computer Science 2017-06-26 Hemanth Venkateswara , Vineeth N. Balasubramanian , Prasanth Lade , Sethuraman Panchanathan

This paper proposes a simple yet effective approach to learn visual features online for improving loop-closure detection and place recognition, based on bag-of-words frameworks. The approach learns a codeword in bag-of-words model from a…

Computer Vision and Pattern Recognition · Computer Science 2016-11-17 Guangcong Zhang , Mason J. Lilly , Patricio A. Vela

Exploring proper way to conduct multi-speech feature fusion for cross-corpus speech emotion recognition is crucial as different speech features could provide complementary cues reflecting human emotion status. While most previous approaches…

Sound · Computer Science 2024-06-14 Xueyu Liu , Jie Lin , Chao Wang

Videos are inherently multimodal. This paper studies the problem of how to fully exploit the abundant multimodal clues for improved video categorization. We introduce a hybrid deep learning framework that integrates useful clues from…

Multimedia · Computer Science 2017-06-15 Yu-Gang Jiang , Zuxuan Wu , Jinhui Tang , Zechao Li , Xiangyang Xue , Shih-Fu Chang

Most popular deep learning based models for action recognition are designed to generate separate predictions within their short temporal windows, which are often aggregated by heuristic means to assign an action label to the full video…

Computer Vision and Pattern Recognition · Computer Science 2017-04-07 Jue Wang , Anoop Cherian , Fatih Porikli , Stephen Gould
‹ Prev 1 4 5 6 7 8 10 Next ›