中文
相关论文

相关论文: Moving Off-the-Grid: Scene-Grounded Video Represen…

200 篇论文

In this paper, a probabilistic space-time representation of complex traffic scenarios is predicted using machine learning algorithms. Such a representation is significant for all active vehicle safety applications especially when performing…

机器学习 · 计算机科学 2025-12-16 Parthasarathy Nadarajan , Michael Botsch , Sebastian Sardina

The appearance of the same object may vary in different scene images due to perspectives and occlusions between objects. Humans can easily identify the same object, even if occlusions exist, by completing the occluded parts based on its…

计算机视觉与模式识别 · 计算机科学 2022-11-28 Tonglin Chen , Bin Li , Zhimeng Shen , Xiangyang Xue

The graph structure is a commonly used data storage mode, and it turns out that the low-dimensional embedded representation of nodes in the graph is extremely useful in various typical tasks, such as node classification, link prediction ,…

社会与信息网络 · 计算机科学 2020-08-03 Xing Li , Wei Wei , Xiangnan Feng , Xue Liu , Zhiming Zheng

This study introduces an efficient and effective method, MeDM, that utilizes pre-trained image Diffusion Models for video-to-video translation with consistent temporal flow. The proposed framework can render videos from scene position…

计算机视觉与模式识别 · 计算机科学 2023-12-21 Ernie Chu , Tzuhsuan Huang , Shuo-Yen Lin , Jun-Cheng Chen

In this paper, we address the novel, highly challenging problem of estimating the layout of a complex urban driving scenario. Given a single color image captured from a driving platform, we aim to predict the bird's-eye view layout of the…

计算机视觉与模式识别 · 计算机科学 2020-02-21 Kaustubh Mani , Swapnil Daga , Shubhika Garg , N. Sai Shankar , Krishna Murthy Jatavallabhula , K. Madhava Krishna

We address the challenge of representation learning from a continuous stream of video as input, in a self-supervised manner. This differs from the standard approaches to video learning where videos are chopped and shuffled during training…

计算机视觉与模式识别 · 计算机科学 2025-04-03 Tengda Han , Dilara Gokay , Joseph Heyward , Chuhan Zhang , Daniel Zoran , Viorica Pătrăucean , João Carreira , Dima Damen , Andrew Zisserman

For robotic agents operating in dynamic environments, learning visual state representations from streaming video observations is essential for sequential decision making. Recent self-supervised learning methods have shown strong…

计算机视觉与模式识别 · 计算机科学 2026-03-26 Seokmin Lee , Yunghee Lee , Byeonghyun Pak , Byeongju Woo

In most interactive image generation tasks, given regions of interest (ROI) by users, the generated results are expected to have adequate diversities in appearance while maintaining correct and reasonable structures in original images. Such…

计算机视觉与模式识别 · 计算机科学 2024-10-28 Jinshu Chen , Qihui Xu , Qi Kang , MengChu Zhou

Event retrieval and recognition in a large corpus of videos necessitates a holistic fixed-size visual representation at the video clip level that is comprehensive, compact, and yet discriminative. It shall comprehensively aggregate…

计算机视觉与模式识别 · 计算机科学 2016-10-12 Zhanning Gao , Gang Hua , Dongqing Zhang , Jianru Xue , Nanning Zheng

Recent advances in unsupervised video object segmentation have highlighted the potential of two-stream architectures that integrate appearance and motion cues. However, fully leveraging these complementary sources of information requires…

计算机视觉与模式识别 · 计算机科学 2026-04-17 Inseok Jeon , Suhwan Cho , Minhyeok Lee , Seunghoon Lee , Minseok Kang , Jungho Lee , Chaewon Park , Donghyeong Kim , Sangyoun Lee

We present a parameterized synthetic dataset called Moving Symbols to support the objective study of video prediction networks. Using several instantiations of the dataset in which variation is explicitly controlled, we highlight issues in…

计算机视觉与模式识别 · 计算机科学 2018-03-23 Ryan Szeto , Simon Stent , German Ros , Jason J. Corso

Recurrent feedback connections in the mammalian visual system have been hypothesized to play a role in synthesizing input in the theoretical framework of analysis by synthesis. The comparison of internally synthesized representation with…

计算机视觉与模式识别 · 计算机科学 2017-05-23 Hao Wang , Xingyu Lin , Yimeng Zhang , Tai Sing Lee

Understanding the shape of a scene from a single color image is a formidable computer vision task. However, most methods aim to predict the geometry of surfaces that are visible to the camera, which is of limited use when planning paths for…

计算机视觉与模式识别 · 计算机科学 2020-04-15 Jamie Watson , Michael Firman , Aron Monszpart , Gabriel J. Brostow

Implicit neural representation, which expresses an image as a continuous function rather than a discrete grid form, is widely used for image processing. Despite its outperforming results, there are still remaining limitations on restoring…

计算机视觉与模式识别 · 计算机科学 2022-09-26 Wonjoon Chang , Dahee Kwon , Bumjin Park

Moving object detection (MOD) is a significant problem in computer vision that has many real world applications. Different categories of methods have been proposed to solve MOD. One of the challenges is to separate moving objects from…

计算机视觉与模式识别 · 计算机科学 2018-08-06 Fateme Bahri , Moein Shakeri , Nilanjan Ray

Large-scale vision-language pre-training has achieved significant performance in multi-modal understanding and generation tasks. However, existing methods often perform poorly on image-text matching tasks that require structured…

计算与语言 · 计算机科学 2023-12-14 Yufeng Huang , Jiji Tang , Zhuo Chen , Rongsheng Zhang , Xinfeng Zhang , Weijie Chen , Zeng Zhao , Zhou Zhao , Tangjie Lv , Zhipeng Hu , Wen Zhang

Video Temporal Grounding (VTG) aims to localize the video segment that corresponds to a natural language query, which requires a comprehensive understanding of complex temporal dynamics. Existing Vision-LMMs typically perceive temporal…

计算机视觉与模式识别 · 计算机科学 2026-03-10 Chaohong Guo , Yihan He , Yongwei Nie , Fei Ma , Xuemiao Xu , Chengjiang Long

We consider the problem of image-to-video translation, where an input image is translated into an output video containing motions of a single object. Recent methods for such problems typically train transformation networks to generate…

计算机视觉与模式识别 · 计算机科学 2018-07-27 Long Zhao , Xi Peng , Yu Tian , Mubbasir Kapadia , Dimitris Metaxas

Online reconstruction of dynamic scenes aims to learn from streaming multi-view inputs under low-latency constraints. The fast training and real-time rendering capabilities of 3D Gaussian Splatting have made on-the-fly reconstruction…

计算机视觉与模式识别 · 计算机科学 2026-03-27 Wonjoon Lee , Sungmin Woo , Donghyeong Kim , Jungho Lee , Sangheon Park , Sangyoun Lee

Multi-view videos are becoming widely used in different fields, but their high resolution and multi-camera shooting raise significant challenges for storage and transmission. In this paper, we propose MV-MGINR, a multi-grid implicit neural…

图像与视频处理 · 电气工程与系统科学 2025-09-23 Qingyue Ling , Zhengxue Cheng , Donghui Feng , Shen Wang , Chen Zhu , Guo Lu , Heming Sun , Jiro Katto , Li Song