English
Related papers

Related papers: Ghrist Barcoded Video Frames. Application in Detec…

200 papers

Extended persistence is a technique from topological data analysis to obtain global multiscale topological information from a graph. This includes information about connected components and cycles that are captured by the so-called…

Machine Learning · Computer Science 2024-06-06 Simon Zhang , Soham Mukherjee , Tamal K. Dey

Well-trained generative neural networks (GNN) are very efficient at compressing visual information for static images in their learned parameters but not as efficient as inter- and intra-prediction for most video content. However, for…

Image and Video Processing · Electrical Eng. & Systems 2020-10-07 Jonah Probell

Graphs provide a powerful framework for modeling complex systems, but their structural variability poses significant challenges for analysis and classification. To address these challenges, we introduce GAUDI (Graph Autoencoder Uncovering…

Machine Learning · Computer Science 2026-02-27 Mirja Granfors , Jesús Pineda , Blanca Zufiria Gerbolés , Joana B. Pereira , Carlo Manzo , Giovanni Volpe

Video Text Spotting (VTS) is a fundamental visual task that aims to predict the trajectories and content of texts in a video. Previous works usually conduct local associations and apply IoU-based distance and complex post-processing…

Computer Vision and Pattern Recognition · Computer Science 2024-01-09 Han Wang , Yanjie Wang , Yang Li , Can Huang

Modern AI-generated videos are photorealistic at the single-frame level, leaving inter-frame dynamics as the main remaining axis for detection. Existing detectors typically handle this temporal evidence in three ways: feeding the full frame…

Computer Vision and Pattern Recognition · Computer Science 2026-05-08 Minsuk Jang , Yujin Yang , Heeseon Kim , Minseok Son , Younghun Kim , Changick Kim

The barcode of a persistence module serves as a complete combinatorial invariant of its isomorphism class. Barcodes are typically extracted by performing changes of basis on a persistence module until the constituent matrices have a special…

Algebraic Topology · Mathematics 2022-07-14 Emile Jacquard , Vidit Nanda , Ulrike Tillmann

A new unified video analytics framework (ER3) is proposed for complex event retrieval, recognition and recounting, based on the proposed video imprint representation, which exploits temporal correlations among image features across video…

Computer Vision and Pattern Recognition · Computer Science 2021-06-08 Zhanning Gao , Le Wang , Nebojsa Jojic , Zhenxing Niu , Nanning Zheng , Gang Hua

The (variational) graph auto-encoder and its variants have been popularly used for representation learning on graph-structured data. While the encoder is often a powerful graph convolutional network, the decoder reconstructs the graph…

Machine Learning · Computer Science 2019-11-27 Han Shi , Haozheng Fan , James T. Kwok

Visual-frame prediction is a pixel-dense prediction task that infers future frames from past frames. Lacking of appearance details, low prediction accuracy and high computational overhead are still major problems with current models or…

Computer Vision and Pattern Recognition · Computer Science 2022-11-16 Chaofan Ling , Junpei Zhong , Weihua Li

Recently, 3D Gaussian Splatting has emerged as a promising approach for modeling 3D scenes using mixtures of Gaussians. The predominant optimization method for these models relies on backpropagating gradients through a differentiable…

Computer Vision and Pattern Recognition · Computer Science 2025-09-24 Toon Van de Maele , Ozan Catal , Alexander Tschantz , Christopher L. Buckley , Tim Verbelen

3D Gaussian Splatting (3DGS) has emerged as a powerful representation due to its efficiency and high-fidelity rendering. 3DGS training requires a known camera pose for each input view, typically obtained by Structure-from-Motion (SfM)…

Computer Vision and Pattern Recognition · Computer Science 2025-05-27 Zhen-Hui Dong , Sheng Ye , Yu-Hui Wen , Nannan Li , Yong-Jin Liu

The encoding of input parameters is one of the fundamental building blocks of neural network algorithms. Its goal is to map the input data to a higher-dimensional space, typically supported by trained feature vectors. The mapping is crucial…

Graphics · Computer Science 2025-07-29 Jakub Bokšanský , Daniel Meister , Carsten Benthin

Recent advances in pretraining general foundation models have significantly improved performance across diverse downstream tasks. While autoregressive (AR) generative models like GPT have revolutionized NLP, most visual generative…

Computer Vision and Pattern Recognition · Computer Science 2025-12-25 Jinghan Li , Yang Jin , Hao Jiang , Yadong Mu , Yang Song , Kun Xu

Recurrent networks of binary neurons are a foundational concept in artificial intelligence. While these networks are traditionally assumed to be fully connected, complex dynamics can emerge when the graph structure is varied. One graph…

Dynamical Systems · Mathematics 2025-08-14 Mirabel Reid , Daniel J. Zhang

This paper introduces the geodesics of triangulated image object shapes. Both rectilinear and curvilinear triangulations of shapes are considered. The triangulation of image object shapes leads to collections of what are known as nerve…

Computational Geometry · Computer Science 2017-08-25 M. Z. Ahmad , J. F. Peters

Topologically ordered phases of matter display a number of unique characteristics, including ground states that can be interpreted as patterns of closed strings. In this paper, we consider the problem of detecting and distinguishing closed…

Statistical Mechanics · Physics 2022-09-16 Dan Sehayek , Roger G. Melko

Recent advances in video generation techniques have given rise to an emerging paradigm of generative video coding for Ultra-Low Bitrate (ULB) scenarios by leveraging powerful generative priors. However, most existing methods are limited by…

Computer Vision and Pattern Recognition · Computer Science 2025-11-12 Zhitao Wang , Hengyu Man , Wenrui Li , Xingtao Wang , Xiaopeng Fan , Debin Zhao

Implicit Neural Representations (INRs) employ neural networks to approximate discrete data as continuous functions. In the context of video data, such models can be utilized to transform the coordinates of pixel locations along with frame…

Computer Vision and Pattern Recognition · Computer Science 2026-02-19 Weronika Smolak-Dyżewska , Dawid Malarz , Kornel Howil , Jan Kaczmarczyk , Marcin Mazur , Przemysław Spurek

This work presents VTok, a unified video tokenization framework that can be used for both generation and understanding tasks. Unlike the leading vision-language systems that tokenize videos through a naive frame-sampling strategy, we…

Computer Vision and Pattern Recognition · Computer Science 2026-02-05 Feng Wang , Yichun Shi , Ceyuan Yang , Qiushan Guo , Jingxiang Sun , Alan Yuille , Peng Wang

$1$-parameter persistent homology, a cornerstone in Topological Data Analysis (TDA), studies the evolution of topological features such as connected components and cycles hidden in data. It has been applied to enhance the representation…

Machine Learning · Computer Science 2023-07-03 Cheng Xin , Soham Mukherjee , Shreyas N. Samaga , Tamal K. Dey