English
Related papers

Related papers: Refined Temporal Pyramidal Compression-and-Amplifi…

200 papers

This paper presents a new method for end-to-end Video Question Answering (VideoQA), aside from the current popularity of using large-scale pre-training with huge feature extractors. We achieve this with a pyramidal multimodal transformer…

Computer Vision and Pattern Recognition · Computer Science 2023-03-07 Min Peng , Chongyang Wang , Yu Shi , Xiang-Dong Zhou

In multivariate time series classification, although current sequence analysis models have excellent classification capabilities, they show significant shortcomings when dealing with long sequence multivariate data, such as prolonged…

Machine Learning · Computer Science 2024-10-29 Enshuo Yan , Huachuan Wang , Weihao Xia

LLM-for-time series (TS) methods typically treat time shallowly, injecting positional or prompt-based cues once at the input of a largely frozen decoder, which limits temporal reasoning as this information degrades through the layers. We…

Machine Learning · Computer Science 2026-02-19 Filippos Bellos , NaveenJohn Premkumar , Yannis Avrithis , Nam H. Nguyen , Jason J. Corso

In this paper a novel method called Extended Two-Dimensional PCA (E2DPCA) is proposed which is an extension to the original 2DPCA. We state that the covariance matrix of 2DPCA is equivalent to the average of the main diagonal of the…

Computer Vision and Pattern Recognition · Computer Science 2010-04-07 Mehran Safayani , Mohammad T. Manzuri-Shalmani , Mahmoud Khademi

This paper proposes a unified framework dubbed Multi-view and Temporal Fusing Transformer (MTF-Transformer) to adaptively handle varying view numbers and video length without camera calibration in 3D Human Pose Estimation (HPE). It consists…

Computer Vision and Pattern Recognition · Computer Science 2022-07-05 Hui Shuai , Lele Wu , Qingshan Liu

Temporal alignment of fine-grained human actions in videos is important for numerous applications in computer vision, robotics, and mixed reality. State-of-the-art methods directly learn image-based embedding space by leveraging powerful…

Computer Vision and Pattern Recognition · Computer Science 2022-04-27 Taein Kwon , Bugra Tekin , Siyu Tang , Marc Pollefeys

Multi-gap Resistive Plate Chamber(MRPC) is a widely used timing detector with a typical time resolution of about 60 ps. This makes MRPC an optimal choice for the time of flight(ToF) system in many large physics experiments. The prior work…

Instrumentation and Detectors · Physics 2019-09-04 Fuyue Wang , Dong Han , Yi Wang , Yancheng Yu , Baohong Guo , Yuanjing Li

Tensor Robust Principal Component Analysis (TRPCA) holds a crucial position in machine learning and computer vision. It aims to recover underlying low-rank structures and to characterize the sparse structures of noise. Current approaches…

Numerical Analysis · Mathematics 2026-01-15 Chao Wang , Huiwen Zheng , Raymond Chan , Youwei Wen

Liquid argon time projection chambers (LArTPCs) provide dense, high-fidelity 3D measurements of particle interactions and underpin current and future neutrino and rare-event experiments. Physics reconstruction typically relies on complex…

High Energy Physics - Experiment · Physics 2025-12-02 Samuel Young , Kazuhiro Terao

Advanced deep Convolutional Neural Networks (CNNs) have shown great success in video-based person Re-Identification (Re-ID). However, they usually focus on the most obvious regions of persons with a limited global representation ability.…

Computer Vision and Pattern Recognition · Computer Science 2023-04-28 Xuehu Liu , Chenyang Yu , Pingping Zhang , Huchuan Lu

We prove under practical assumptions that Rotary Positional Embedding (RoPE) introduces an intrinsic distance-dependent bias in attention scores that limits RoPE's ability to model long-context. RoPE extension methods may alleviate this…

Computation and Language · Computer Science 2026-05-12 Yu Wang , Sheng Shen , Rémi Munos , Hongyuan Zhan , Yuandong Tian

Temporal 3D human pose estimation from monocular videos is a challenging task in human-centered computer vision due to the depth ambiguity of 2D-to-3D lifting. To improve accuracy and address occlusion issues, inertial sensor has been…

Computer Vision and Pattern Recognition · Computer Science 2024-04-30 Yiming Bao , Xu Zhao , Dahong Qian

Understanding causal event relationships and achieving fine-grained temporal grounding in videos remain challenging for vision-language models. Existing methods either compress video tokens to reduce temporal resolution, or treat videos as…

Recently, fully-transformer architectures have replaced the defacto convolutional architecture for the 3D human pose estimation task. In this paper we propose \textbf{\textit{ConvFormer}}, a novel convolutional transformer that leverages a…

Computer Vision and Pattern Recognition · Computer Science 2023-04-06 Alec Diaz-Arias , Dmitriy Shin

Reconstructing 3D human pose and shape from monocular videos is a well-studied but challenging problem. Common challenges include occlusions, the inherent ambiguities in the 2D to 3D mapping and the computational complexity of video…

Computer Vision and Pattern Recognition · Computer Science 2023-05-02 Nikolaos Vasilikopoulos , Nikos Kolotouros , Aggeliki Tsoli , Antonis Argyros

This paper introduces a novel Pre-trained Spatial Temporal Many-to-One (P-STMO) model for 2D-to-3D human pose estimation task. To reduce the difficulty of capturing spatial and temporal information, we divide this task into two stages:…

Computer Vision and Pattern Recognition · Computer Science 2022-08-01 Wenkang Shan , Zhenhua Liu , Xinfeng Zhang , Shanshe Wang , Siwei Ma , Wen Gao

Head-related transfer function (HRTF) plays an important role in the construction of 3D auditory display. This paper presents an individual HRTF modeling method using deep neural networks based on spatial principal component analysis. The…

Audio and Speech Processing · Electrical Eng. & Systems 2020-02-07 Mengfan Zhang , Zhongshu Ge , Tiejun Liu , Xihong Wu , Tianshu Qu

Temporal Action Detection (TAD) in untrimmed videos poses significant challenges, particularly for Activities of Daily Living (ADL) requiring models to (1) process long-duration videos, (2) capture temporal variations in actions, and (3)…

Computer Vision and Pattern Recognition · Computer Science 2025-12-18 Arkaprava Sinha , Monish Soundar Raj , Pu Wang , Ahmed Helmy , Hieu Le , Srijan Das

Vision-Language-Action models (VLA) have demonstrated remarkable capabilities and promising potential in solving complex robotic manipulation tasks. However, their substantial parameter sizes and high inference latency pose significant…

Robotics · Computer Science 2025-06-24 Yuxuan Chen , Xiao Li

Despite the recent success of single image-based 3D human pose and shape estimation methods, recovering temporally consistent and smooth 3D human motion from a video is still challenging. Several video-based methods have been proposed;…

Computer Vision and Pattern Recognition · Computer Science 2021-04-28 Hongsuk Choi , Gyeongsik Moon , Ju Yong Chang , Kyoung Mu Lee