English
Related papers

Related papers: TEAM-Net: Multi-modal Learning for Video Action Re…

200 papers

Video compression technology is essential for transmitting and storing videos. Many video compression methods reduce information in videos by removing high-frequency components and utilizing similarities between frames. Alternatively, the…

Computer Vision and Pattern Recognition · Computer Science 2024-07-09 Taiga Hayami , Hiroshi Watanabe

Group activity detection in multi-person scenes is challenging due to complex human interactions, occlusions, and variations in appearance over time. This work presents a computer vision based framework for group activity recognition and…

Computer Vision and Pattern Recognition · Computer Science 2025-11-18 Narthana Sivalingam , Santhirarajah Sivasthigan , Thamayanthi Mahendranathan , G. M. R. I. Godaliyadda , M. P. B. Ekanayake , H. M. V. R. Herath

While traditional methods relies on depth sensors, the current trend leans towards utilizing cost-effective RGB images, despite their absence of depth cues. This paper introduces an interesting approach to detect grasping pose from a single…

Computer Vision and Pattern Recognition · Computer Science 2023-10-31 Zhaocong Li

One key challenge to learning-based video compression is that motion predictive coding, a very effective tool for video compression, can hardly be trained into a neural network. In this paper we propose the concept of PixelMotionCNN (PMCNN)…

Multimedia · Computer Science 2019-01-15 Zhibo Chen , Tianyu He , Xin Jin , Feng Wu

Visual-semantic embedding aims to find a shared latent space where related visual and textual instances are close to each other. Most current methods learn injective embedding functions that map an instance to a single point in the shared…

Computer Vision and Pattern Recognition · Computer Science 2019-07-18 Yale Song , Mohammad Soleymani

The majority of learning-based semantic segmentation methods are optimized for daytime scenarios and favorable lighting conditions. Real-world driving scenarios, however, entail adverse environmental conditions such as nighttime…

Computer Vision and Pattern Recognition · Computer Science 2020-03-11 Johan Vertens , Jannik Zürn , Wolfram Burgard

In this paper, a mode selection network (ModeNet) is proposed to enhance deep learning-based video compression. Inspired by traditional video coding, ModeNet purpose is to enable competition among several coding modes. The proposed ModeNet…

Neural and Evolutionary Computing · Computer Science 2020-08-03 Théo Ladune , Pierrick Philippe , Wassim Hamidouche , Lu Zhang , Olivier Déforges

Human action recognition is regarded as a key cornerstone in domains such as surveillance or video understanding. Despite recent progress in the development of end-to-end solutions for video-based action recognition, achieving…

Computer Vision and Pattern Recognition · Computer Science 2020-08-04 Jiawei Chen , Jenson Hsiao , Chiu Man Ho

Event cameras provide several unique advantages over standard frame-based sensors, including high temporal resolution, low latency, and robustness to extreme lighting. However, existing learning-based approaches for event processing are…

Computer Vision and Pattern Recognition · Computer Science 2026-05-05 Vincenzo Polizzi , David B. Lindell , Jonathan Kelly

Unsupervised video object segmentation aims to detect the most salient object in a video without any external guidance regarding the object. Salient objects often exhibit distinctive movements compared to the background, and recent methods…

Computer Vision and Pattern Recognition · Computer Science 2025-04-28 Suhwan Cho , Minhyeok Lee , Jungho Lee , MyeongAh Cho , Seungwook Park , Jaeyeob Kim , Hyunsung Jang , Sangyoun Lee

The multi-modal salient object detection model based on RGB-D information has better robustness in the real world. However, it remains nontrivial to better adaptively balance effective multi-modal information in the feature fusion phase. In…

Computer Vision and Pattern Recognition · Computer Science 2022-02-09 Jinchao Zhu , Xiaoyu Zhang , Xian Fang , Feng Dong , Qiu Yu

RGB-Thermal Video Object Detection (RGBT VOD) can address the limitation of traditional RGB-based VOD in challenging lighting conditions, making it more practical and effective in many applications. However, similar to most RGBT fusion…

Computer Vision and Pattern Recognition · Computer Science 2025-04-17 Qishun Wang , Zhengzheng Tu , Chenglong Li , Bo Jiang

Heterogeneous data modalities can provide complementary cues for several tasks, usually leading to more robust algorithms and better performance. However, while training data can be accurately collected to include a variety of sensory…

Computer Vision and Pattern Recognition · Computer Science 2019-07-29 Nuno C. Garcia , Pietro Morerio , Vittorio Murino

Understanding continuous video streams plays a fundamental role in real-time applications including embodied AI and autonomous driving. Unlike offline video understanding, streaming video understanding requires the ability to process video…

Computer Vision and Pattern Recognition · Computer Science 2025-07-23 Yibin Yan , Jilan Xu , Shangzhe Di , Yikun Liu , Yudi Shi , Qirui Chen , Zeqian Li , Yifei Huang , Weidi Xie

Many self-supervised learning methods are pre-trained on the well-curated ImageNet-1K dataset. In this work, given the excellent scalability of web data, we consider self-supervised pre-training on noisy web sourced image-text paired data.…

Computer Vision and Pattern Recognition · Computer Science 2024-08-06 Bingchen Zhao , Quan Cui , Hao Wu , Osamu Yoshie , Cheng Yang , Oisin Mac Aodha

Neural fields, also known as coordinate-based or implicit neural representations, have shown a remarkable capability of representing, generating, and manipulating various forms of signals. For video representations, however, mapping…

Computer Vision and Pattern Recognition · Computer Science 2023-08-08 Joo Chan Lee , Daniel Rho , Jong Hwan Ko , Eunbyung Park

Video captioning is a challenging task since it requires generating sentences describing various diverse and complex videos. Existing video captioning models lack adequate visual representation due to the neglect of the existence of gaps…

Computer Vision and Pattern Recognition · Computer Science 2021-10-14 Mingkang Tang , Zhanyu Wang , Zhenhua Liu , Fengyun Rao , Dian Li , Xiu Li

Human actions comprise of joint motion of articulated body parts or `gestures'. Human skeleton is intuitively represented as a sparse graph with joints as nodes and natural connections between them as edges. Graph convolutional networks…

Computer Vision and Pattern Recognition · Computer Science 2018-09-14 Kalpit Thakkar , P J Narayanan

We address the problem of temporal localization of repetitive activities in a video, i.e., the problem of identifying all segments of a video that contain some sort of repetitive or periodic motion. To do so, the proposed method represents…

Computer Vision and Pattern Recognition · Computer Science 2019-10-15 Giorgos Karvounas , Iason Oikonomidis , Antonis Argyros

There has been huge progress on video action recognition in recent years. However, many works focus on tweaking existing 2D backbones due to the reliance of ImageNet pretraining, which restrains the models from achieving higher efficiency…

Computer Vision and Pattern Recognition · Computer Science 2025-03-05 Zhe Wang , Xulei Yang
‹ Prev 1 8 9 10 Next ›