中文
相关论文

相关论文: Seeing without Pixels: Perception from Camera Traj…

200 篇论文

This paper introduces a novel method for self-supervised video representation learning via feature prediction. In contrast to the previous methods that focus on future feature prediction, we argue that a supervisory signal arising from…

计算机视觉与模式识别 · 计算机科学 2020-11-13 Nadine Behrmann , Juergen Gall , Mehdi Noroozi

In the absence of a mechanical stabilizer, the camera undergoes inevitable rotational dynamics during capturing, which induces perspective-based blur especially under long-exposure scenarios. From an optical standpoint, perspective-based…

计算机视觉与模式识别 · 计算机科学 2025-12-10 Tianchen Qiu , Qirun Zhang , Jiajian He , Zhengyue Zhuge , Jiahui Xu , Yueting Chen

We explain theoretically how to reconstruct the 3D scene from successive frames in order to see the video in 3D. To do this, features, associated to moving rigid objects in 3D, are extracted in frames and matched. The vanishing point…

计算机视觉与模式识别 · 计算机科学 2020-07-30 Slimane Larabi

Distilling analytical models from data has the potential to advance our understanding and prediction of nonlinear dynamics. Although discovery of governing equations based on observed system states (e.g., trajectory time series) has…

机器学习 · 计算机科学 2021-06-10 Lele Luan , Yang Liu , Hao Sun

In this paper we propose a novel approach for detecting and tracking objects in videos with variable background i.e. videos captured by moving cameras without any additional sensor. In a video captured by a moving camera, both the…

计算机视觉与模式识别 · 计算机科学 2017-05-09 Kumar S. Ray , Vijayan K. Asari , Soma Chakraborty

Motion segmentation in dynamic scenes is highly challenging, as conventional methods heavily rely on estimating camera poses and point correspondences from inherently noisy motion cues. Existing statistical inference or iterative…

计算机视觉与模式识别 · 计算机科学 2026-02-26 Xiankang He , Peile Lin , Ying Cui , Dongyan Guo , Chunhua Shen , Xiaoqin Zhang

We present a method for unsupervised learning of equations of motion for objects in raw and optionally distorted unlabeled video. We first train an autoencoder that maps each video frame into a low-dimensional latent space where the laws of…

计算机视觉与模式识别 · 计算机科学 2021-04-28 Silviu-Marian Udrescu , Max Tegmark

A natural approach to generative modeling of videos is to represent them as a composition of moving objects. Recent works model a set of 2D sprites over a slowly-varying background, but without considering the underlying 3D scene that gives…

计算机视觉与模式识别 · 计算机科学 2021-03-26 Paul Henderson , Christoph H. Lampert

This work presents an unsupervised learning based approach to the ubiquitous computer vision problem of image matching. We start from the insight that the problem of frame-interpolation implicitly solves for inter-frame correspondences.…

计算机视觉与模式识别 · 计算机科学 2016-03-30 Gucan Long , Laurent Kneip , Jose M. Alvarez , Hongdong Li

Camera calibration is integral to robotics and computer vision algorithms that seek to infer geometric properties of the scene from visual input streams. In practice, calibration is a laborious procedure requiring specialized data…

计算机视觉与模式识别 · 计算机科学 2022-03-03 Jiading Fang , Igor Vasiljevic , Vitor Guizilini , Rares Ambrus , Greg Shakhnarovich , Adrien Gaidon , Matthew R. Walter

We introduce Perception Encoder (PE), a state-of-the-art vision encoder for image and video understanding trained via simple vision-language learning. Traditionally, vision encoders have relied on a variety of pretraining objectives, each…

Recent advancements in video generation have been greatly driven by video diffusion models, with camera motion control emerging as a crucial challenge in creating view-customized visual content. This paper introduces trajectory attention, a…

计算机视觉与模式识别 · 计算机科学 2024-12-02 Zeqi Xiao , Wenqi Ouyang , Yifan Zhou , Shuai Yang , Lei Yang , Jianlou Si , Xingang Pan

Human actions are comprised of a sequence of poses. This makes videos of humans a rich and dense source of human poses. We propose an unsupervised method to learn pose features from videos that exploits a signal which is complementary to…

计算机视觉与模式识别 · 计算机科学 2016-09-20 Senthil Purushwalkam , Abhinav Gupta

In this paper, we propose to learn temporal embeddings of video frames for complex video analysis. Large quantities of unlabeled video data can be easily obtained from the Internet. These videos possess the implicit weak label that they are…

计算机视觉与模式识别 · 计算机科学 2015-05-05 Vignesh Ramanathan , Kevin Tang , Greg Mori , Li Fei-Fei

In this paper, we introduce a novel method to capture visual trajectories for navigating an indoor robot in dynamic settings using streaming image data. First, an image processing pipeline is proposed to accurately segment trajectories from…

机器人学 · 计算机科学 2020-01-13 Aditya Rajguru , Christopher Collander , William J. Beksi

Accurate observation of dynamic environments traditionally relies on synthesizing raw, signal-level information from multiple distributed sensors. This work investigates an alternative approach: performing geospatial inference using only…

计算机视觉与模式识别 · 计算机科学 2026-03-31 Sadik Yagiz Yetim , Gaofeng Dong , Isaac-Neil Zanoria , Ronit Barman , Maggie Wigness , Tarek Abdelzaher , Mani Srivastava , Suhas Diggavi

Diffusion models have made significant strides in image generation, mastering tasks such as unconditional image synthesis, text-image translation, and image-to-image conversions. However, their capability falls short in the realm of video…

计算机视觉与模式识别 · 计算机科学 2024-12-10 Gaurav Shrivastava , Abhinav Shrivastava

The objective of this paper is self-supervised learning of feature embeddings that are suitable for matching correspondences along the videos, which we term correspondence flow. By leveraging the natural spatial-temporal coherence in…

计算机视觉与模式识别 · 计算机科学 2019-07-30 Zihang Lai , Weidi Xie

Standard video encoders developed for conventional narrow field-of-view video are widely applied to 360{\deg} video as well, with reasonable results. However, while this approach commits arbitrarily to a projection of the spherical frames,…

计算机视觉与模式识别 · 计算机科学 2017-12-13 Yu-Chuan Su , Kristen Grauman

We present an algorithm to estimate depth in dynamic video scenes. We propose to learn and infer depth in videos from appearance, motion, occlusion boundaries, and geometric context of the scene. Using our method, depth can be estimated…

计算机视觉与模式识别 · 计算机科学 2015-10-27 S. Hussain Raza , Omar Javed , Aveek Das , Harpreet Sawhney , Hui Cheng , Irfan Essa