English
Related papers

Related papers: A Heterogeneous Two-Stream Framework for Video Act…

200 papers

Foundation models (FMs) are large neural networks trained on broad datasets, excelling in downstream tasks with minimal fine-tuning. Human activity recognition in video has advanced with FMs, driven by competition among different…

Computer Vision and Pattern Recognition · Computer Science 2024-07-23 Thinesh Thiyakesan Ponbagavathi , Kunyu Peng , Alina Roitberg

Weakly supervised video anomaly detection (WS-VAD) is a crucial area in computer vision for developing intelligent surveillance systems. This system uses three feature streams: RGB video, optical flow, and audio signals, where each stream…

Computer Vision and Pattern Recognition · Computer Science 2025-04-07 Yuta Kaneko , Abu Saleh Musa Miah , Najmul Hassan , Hyoun-Sup Lee , Si-Woong Jang , Jungpil Shin

Unsupervised video object segmentation (UVOS) aims at detecting the primary objects in a given video sequence without any human interposing. Most existing methods rely on two-stream architectures that separately encode the appearance and…

Computer Vision and Pattern Recognition · Computer Science 2023-12-01 Lingyi Hong , Wei Zhang , Shuyong Gao , Hong Lu , WenQiang Zhang

Occlusions between consecutive frames have long posed a significant challenge in optical flow estimation. The inherent ambiguity introduced by occlusions directly violates the brightness constancy constraint and considerably hinders…

Computer Vision and Pattern Recognition · Computer Science 2023-11-30 Shangkun Sun , Jiaming Liu , Thomas H. Li , Huaxia Li , Guoqing Liu , Wei Gao

Accurate fovea localization is essential for analyzing retinal diseases to prevent irreversible vision loss. While current deep learning-based methods outperform traditional ones, they still face challenges such as the lack of local…

Computer Vision and Pattern Recognition · Computer Science 2024-10-11 Sifan Song , Jinfeng Wang , Zilong Wang , Hongxing Wang , Jionglong Su , Xiaowei Ding , Kang Dang

Motion estimation is one of the core challenges in computer vision. With traditional dual-frame approaches, occlusions and out-of-view motions are a limiting factor, especially in the context of environmental perception for vehicles due to…

Computer Vision and Pattern Recognition · Computer Science 2020-11-05 René Schuster , Christian Unger , Didier Stricker

Distributed radar sensors enable robust human activity recognition. However, scaling the number of coordinated nodes introduces challenges in feature extraction from large datasets, and transparent data fusion. We propose an end-to-end…

Signal Processing · Electrical Eng. & Systems 2026-01-07 Mina Shahbazifar , Zolfa Zeinalpour-Yazdi , Matthias Hollick , Arash Asadi , Vahid Jamali

Existing RGB-Event visual object tracking approaches primarily rely on conventional feature-level fusion, failing to fully exploit the unique advantages of event cameras. In particular, the high dynamic range and motion-sensitive nature of…

Computer Vision and Pattern Recognition · Computer Science 2026-01-06 Shiao Wang , Xiao Wang , Haonan Zhao , Jiarui Xu , Bo Jiang , Lin Zhu , Xin Zhao , Yonghong Tian , Jin Tang

This paper presents a novel hybrid representation learning framework for streaming data, where an image frame in a video is modeled by an ensemble of two distinct deep neural networks; one is a low-bit quantized network and the other is a…

Computer Vision and Pattern Recognition · Computer Science 2022-05-24 Ilchae Jung , Minji Kim , Eunhyeok Park , Bohyung Han

We bring together ideas from recent work on feature design for egocentric action recognition under one framework by exploring the use of deep convolutional neural networks (CNN). Recent work has shown that features such as hand appearance,…

Computer Vision and Pattern Recognition · Computer Science 2016-05-13 Minghuang Ma , Haoqi Fan , Kris M. Kitani

Data-fusion networks have shown significant promise for RGB-thermal scene parsing. However, the majority of existing studies have relied on symmetric duplex encoders for heterogeneous feature extraction and fusion, paying inadequate…

Computer Vision and Pattern Recognition · Computer Science 2026-01-07 Jiahang Li , Peng Yun , Yang Xu , Ye Zhang , Mingjian Sun , Qijun Chen , Ilin Alexander , Rui Fan

Robust gait recognition requires highly discriminative representations, which are closely tied to input modalities. While binary silhouettes and skeletons have dominated recent literature, these 2D representations fall short of capturing…

Computer Vision and Pattern Recognition · Computer Science 2025-08-06 Xinzhu Li , Juepeng Zheng , Yikun Chen , Xudong Mao , Guanghui Yue , Wei Zhou , Chenlei Lv , Ruomei Wang , Fan Zhou , Baoquan Zhao

Motion recognition is a promising direction in computer vision, but the training of video classification models is much harder than images due to insufficient data and considerable parameters. To get around this, some works strive to…

Computer Vision and Pattern Recognition · Computer Science 2023-06-13 Benjia Zhou , Pichao Wang , Jun Wan , Yanyan Liang , Fan Wang

Networks of interconnected resistors, springs and beams, or pores are standard models of studying scalar and vector transport processes in heterogeneous materials and media, such as fluid flow in porous media, and conduction, deformations,…

Computational Physics · Physics 2019-08-12 Hassan Dashtian , Muhammad Sahimi

RGB-Thermal Salient Object Detection aims to pinpoint prominent objects within aligned pairs of visible and thermal infrared images. Traditional encoder-decoder architectures, while designed for cross-modality feature interactions, may not…

Computer Vision and Pattern Recognition · Computer Science 2024-12-03 Hao Tang , Zechao Li , Dong Zhang , Shengfeng He , Jinhui Tang

Unsupervised optical flow methods typically lack reliable uncertainty estimation, limiting their robustness and interpretability. We propose U$^{2}$Flow, the first recurrent unsupervised framework that jointly estimates optical flow and…

Computer Vision and Pattern Recognition · Computer Science 2026-04-14 Xunpei Sun , Wenwei Lin , Yi Chang , Gang Chen

Multi-sensor fusion is central to robust robotic perception, yet most existing systems operate under static sensor configurations, collecting all modalities at fixed rates and fidelity regardless of their situational utility. This rigidity…

Robotics · Computer Science 2026-02-12 Yanchen Liu , Yuang Fan , Minghui Zhao , Xiaofan Jiang

Current optical flow methods exploit the stable appearance of frame (or RGB) data to establish robust correspondences across time. Event cameras, on the other hand, provide high-temporal-resolution motion cues and excel in challenging…

Computer Vision and Pattern Recognition · Computer Science 2025-08-20 Qianang Zhou , Junhui Hou , Meiyi Yang , Yongjian Deng , Youfu Li , Junlin Xiong

In this paper, several variants of two-stream architectures for temporal action proposal generation in long, untrimmed videos are presented. Inspired by the recent advances in the field of human action recognition utilizing 3D convolutions…

Computer Vision and Pattern Recognition · Computer Science 2019-03-15 Patrick Schlosser , David Münch , Michael Arens

The temporal component of videos provides an important clue for activity recognition, as a number of activities can be reliably recognized based on the motion information. In view of that, this work proposes a novel temporal stream for…

Computer Vision and Pattern Recognition · Computer Science 2017-08-23 Carlos Caetano , Victor H. C. de Melo , Jefersson A. dos Santos , William Robson Schwartz