English
Related papers

Related papers: Volume-Based Space-Time Cube for Large-Scale Conti…

200 papers

Large-scale video-language pretraining enables strong generalization across multimodal tasks but often incurs prohibitive computational costs. Although recent advances in masked visual modeling help mitigate this issue, they still suffer…

Computer Vision and Pattern Recognition · Computer Science 2026-03-25 Weijun Zhuang , Yuqing Huang , Weikang Meng , Xin Li , Ming Liu , Xiaopeng Hong , Yaowei Wang , Wangmeng Zuo

Rapid Serial Visual Presentation (RSVP) is a paradigm that supports the application of cortically coupled computer vision to rapid image search. In RSVP, images are presented to participants in a rapid serial sequence which can evoke…

Image and Video Processing · Electrical Eng. & Systems 2019-01-16 Zhengwei Wang , Graham Healy , Alan F. Smeaton , Tomas E. Ward

Deep learning models have enjoyed great success for image related computer vision tasks like image classification and object detection. For video related tasks like human action recognition, however, the advancements are not as significant…

Computer Vision and Pattern Recognition · Computer Science 2018-09-12 Xiaolin Song , Cuiling Lan , Wenjun Zeng , Junliang Xing , Jingyu Yang , Xiaoyan Sun

Surgical video understanding is a crucial prerequisite for advancing Computer-Assisted Surgery. While vision-language models (VLMs) have recently been applied to the surgical domain, existing surgical vision-language datasets lack in…

Computer Vision and Pattern Recognition · Computer Science 2026-04-02 Lennart Maack , Alexander Schlaefer

Recent advancements in video super-resolution (VSR) models have demonstrated impressive results in enhancing low-resolution videos. However, due to limitations in adequately controlling the generation process, achieving high fidelity…

Computer Vision and Pattern Recognition · Computer Science 2025-08-04 Yiwen Wang , Xinning Chai , Yuhong Zhang , Zhengxue Cheng , Jun Zhao , Rong Xie , Li Song

Numerous studies have investigated the pivotal role of reliable 3D volume representation in scene perception tasks, such as multi-view stereo (MVS) and semantic scene completion (SSC). They typically construct 3D probability volumes…

Computer Vision and Pattern Recognition · Computer Science 2024-01-30 Bohan Li , Yasheng Sun , Jingxin Dong , Zheng Zhu , Jinming Liu , Xin Jin , Wenjun Zeng

This paper proposes a spatiotemporal (ST) fusion framework robust against diverse noise for satellite images, named Temporally-Similar Structure-Aware ST fusion (TSSTF). ST fusion is a promising approach to address the trade-off between the…

Signal Processing · Electrical Eng. & Systems 2026-01-30 Ryosuke Isono , Shunsuke Ono

High level understanding of sequential visual input is important for safe and stable autonomy, especially in localization and object detection. While traditional object classification and tracking approaches are specifically designed to…

Computer Vision and Pattern Recognition · Computer Science 2017-07-25 Mo Shan , Nikolay Atanasov

Data-driven soft sensors have been widely applied in complex industrial processes. However, the interpretable spatio-temporal features extraction by soft sensors remains a challenge. In this light, this work introduces a novel method termed…

Systems and Control · Electrical Eng. & Systems 2025-11-04 Qianchao Wang , Peng Sha , Leena Heistrene , Yuxuan Ding , Yaping Du

This paper aims to address the challenge of reconstructing long volumetric videos from multi-view RGB videos. Recent dynamic view synthesis methods leverage powerful 4D representations, like feature grids or point cloud sequences, to…

Computer Vision and Pattern Recognition · Computer Science 2024-12-13 Zhen Xu , Yinghao Xu , Zhiyuan Yu , Sida Peng , Jiaming Sun , Hujun Bao , Xiaowei Zhou

Cellular identity and function are linked to both their intrinsic genomic makeup and extrinsic spatial context within the tissue microenvironment. Spatial transcriptomics (ST) offers an unprecedented opportunity to study this, providing in…

Machine Learning · Computer Science 2026-02-16 Rui Yan , Xiaohan Xing , Xun Wang , Zixia Zhou , Md Tauhidul Islam , Lei Xing

Large Language Models (LLMs) have showcased impressive capabilities in text comprehension and generation, prompting research efforts towards video LLMs to facilitate human-AI interaction at the video level. However, how to effectively…

Computer Vision and Pattern Recognition · Computer Science 2024-04-02 Ruyang Liu , Chen Li , Haoran Tang , Yixiao Ge , Ying Shan , Ge Li

Generating video descriptions automatically is a challenging task that involves a complex interplay between spatio-temporal visual features and language models. Given that videos consist of spatial (frame-level) features and their temporal…

Computer Vision and Pattern Recognition · Computer Science 2020-01-20 Anoop Cherian , Jue Wang , Chiori Hori , Tim K. Marks

Vision-Language Models (VLMs) have become central to autonomous driving systems, yet their deployment is severely bottlenecked by the massive computational overhead of multi-view camera and multi-frame video input. Existing token pruning…

Computer Vision and Pattern Recognition · Computer Science 2026-04-22 Lin Sha , Haiyun Guo , Tao Wang , Cong Zhang , Min Huang , Jinqiao Wang , Qinghai Miao

Taxi arrival time prediction is essential for building intelligent transportation systems. Traditional prediction methods mainly rely on extracting features from traffic maps, which cannot model complex situations and nonlinear spatial and…

Machine Learning · Computer Science 2021-10-01 ZiChuan Liu , Zhaoyang Wu , Meng Wang , Rui Zhang

In this paper, we propose a spatio-temporal contextual network, STC-Flow, for optical flow estimation. Unlike previous optical flow estimation approaches with local pyramid feature extraction and multi-level correlation, we propose a…

Computer Vision and Pattern Recognition · Computer Science 2020-11-04 Xiaolin Song , Yuyang Zhao , Jingyu Yang

Video-based gaze estimation methods aim to capture the inherently temporal dynamics of human eye gaze from multiple image frames. However, since models must capture both spatial and temporal relationships, performance is limited by the…

Computer Vision and Pattern Recognition · Computer Science 2025-12-22 Alexandre Personnic , Mihai Bâce

As big spatial data becomes increasingly prevalent, classical spatiotemporal (ST) methods often do not scale well. While methods have been developed to account for high-dimensional spatial objects, the setting where there are exceedingly…

Applications · Statistics 2019-08-27 Samuel I. Berchuck , Felipe A. Medeiros , Sayan Mukherjee

Discovering dominant patterns and exploring dynamic behaviors especially critical state transitions and tipping points in high-dimensional time-series data are challenging tasks in study of real-world complex systems, which demand…

Machine Learning · Statistics 2025-01-23 Pei Chen , Yaofang Suo , Rui Liu , Luonan Chen

The rapid development of facial manipulation techniques has aroused public concerns in recent years. Following the success of deep learning, existing methods always formulate DeepFake video detection as a binary classification problem and…

Computer Vision and Pattern Recognition · Computer Science 2021-10-12 Zhihao Gu , Yang Chen , Taiping Yao , Shouhong Ding , Jilin Li , Feiyue Huang , Lizhuang Ma
‹ Prev 1 3 4 5 6 7 10 Next ›