中文
相关论文

相关论文: Sequential Attend, Infer, Repeat: Generative Model…

200 篇论文

Many visual phenomena suggest that humans use top-down generative or reconstructive processes to create visual percepts (e.g., imagery, object completion, pareidolia), but little is known about the role reconstruction plays in robust object…

计算机视觉与模式识别 · 计算机科学 2023-02-09 Seoyoung Ahn , Hossein Adeli , Gregory J. Zelinsky

In this paper, we propose a scene-level inverse rendering framework that uses multi-view images to decompose the scene into geometry, SVBRDF, and 3D spatially-varying lighting. While multi-view images have been widely used for object-level…

计算机视觉与模式识别 · 计算机科学 2025-03-07 JunYong Choi , SeokYeong Lee , Haesol Park , Seung-Won Jung , Ig-Jae Kim , Junghyun Cho

We present STaR, a novel method that performs Self-supervised Tracking and Reconstruction of dynamic scenes with rigid motion from multi-view RGB videos without any manual annotation. Recent work has shown that neural networks are…

计算机视觉与模式识别 · 计算机科学 2021-01-06 Wentao Yuan , Zhaoyang Lv , Tanner Schmidt , Steven Lovegrove

Extensions of Neural Radiance Fields (NeRFs) to model dynamic scenes have enabled their near photo-realistic, free-viewpoint rendering. Although these methods have shown some potential in creating immersive experiences, two drawbacks limit…

计算机视觉与模式识别 · 计算机科学 2024-06-07 Xinhang Liu , Yu-Wing Tai , Chi-Keung Tang , Pedro Miraldo , Suhas Lohit , Moitreya Chatterjee

A grand challenge in machine learning is the development of computational algorithms that match or outperform humans in perceptual inference tasks that are complicated by nuisance variation. For instance, visual object recognition involves…

机器学习 · 统计学 2015-04-03 Ankit B. Patel , Tan Nguyen , Richard G. Baraniuk

In this paper we start with a simple question, how is it possible that humans can recognize different movements over skin with only a prior visual experience of them? Or in general, what is the representation of spatial sequences that are…

人工智能 · 计算机科学 2023-11-14 Viacheslav M. Osaulenko

Autoregressive models have recently shown great promise in visual generation by leveraging discrete token sequences akin to language modeling. However, existing approaches often suffer from inefficiency, either due to token-by-token…

计算机视觉与模式识别 · 计算机科学 2025-11-20 Ruiqing Yang , Kaixin Zhang , Zheng Zhang , Shan You , Tao Huang

We explore recurrent encoder multi-decoder neural network architectures for semi-supervised sequence classification and reconstruction. We find that the use of multiple reconstruction modules helps models generalize in a classification task…

计算机视觉与模式识别 · 计算机科学 2018-07-12 Félix G. Harvey , Julien Roy , David Kanaa , Christopher Pal

Fine-grained image recognition has been a hot research topic in computer vision due to its various applications. The-state-of-the-art is the part/region-based approaches that first localize discriminative parts/regions, and then learn their…

计算机视觉与模式识别 · 计算机科学 2019-08-07 Peng Zhang , Xinyu Zhu , Zhanzhan Cheng , Shuigeng Zhou , Yi Niu

Grasp planning and estimation have been a longstanding research problem in robotics, with two main approaches to find graspable poses on the objects: 1) geometric approach, which relies on 3D models of objects and the gripper to estimate…

机器人学 · 计算机科学 2025-04-11 Xun Tu , Karthik Desingh

We propose SLARM, a feed-forward model that unifies dynamic scene reconstruction, semantic understanding, and real-time streaming inference. SLARM captures complex, non-uniform motion through higher-order motion modeling, trained solely on…

计算机视觉与模式识别 · 计算机科学 2026-03-27 Zhicheng Qiu , Jiarui Meng , Tong-an Luo , Yican Huang , Xuan Feng , Xuanfu Li , ZHan Xu

This paper proposes a method for performing continual learning of predictive models that facilitate the inference of future frames in video sequences. For a first given experience, an initial Variational Autoencoder, together with a set of…

计算机视觉与模式识别 · 计算机科学 2020-06-04 Damian Campo , Giulia Slavic , Mohamad Baydoun , Lucio Marcenaro , Carlo Regazzoni

Gait recognition is a rapidly advancing vision technique for person identification from a distance. Prior studies predominantly employed relatively shallow networks to extract subtle gait features, achieving impressive successes in…

计算机视觉与模式识别 · 计算机科学 2024-01-11 Chao Fan , Saihui Hou , Yongzhen Huang , Shiqi Yu

Despite the astonishing progress in generative AI, 4D dynamic object generation remains an open challenge. With limited high-quality training data and heavy computing requirements, the combination of hallucinating unseen geometry together…

计算机视觉与模式识别 · 计算机科学 2025-04-18 Lu Sang , Zehranaz Canfes , Dongliang Cao , Riccardo Marin , Florian Bernard , Daniel Cremers

Sequential dense retrieval models utilize advanced sequence learning techniques to compute item and user representations, which are then used to rank relevant items for a user through inner product computation between the user and all item…

We propose a new method for inferring the governing stochastic ordinary differential equations (SODEs) by observing particle ensembles at discrete and sparse time instants, i.e., multiple "snapshots". Particle coordinates at a single time…

机器学习 · 计算机科学 2021-03-23 Liu Yang , Constantinos Daskalakis , George Em Karniadakis

The rapid progression of generative AI (GenAI) technologies has heightened concerns regarding the misuse of AI-generated imagery. To address this issue, robust detection methods have emerged as particularly compelling, especially in…

图形学 · 计算机科学 2025-04-07 Hongfei Cai , Chi Liu , Sheng Shen , Youyang Qu , Peng Gui

Prediction with high accuracy is essential for various applications such as autonomous driving. Existing prediction models are easily prone to errors in real-world settings where observations (e.g. human poses and locations) are often…

计算机视觉与模式识别 · 计算机科学 2021-03-29 Manh Huynh , Gita Alaghband

Learning geometry, motion, and appearance priors of object classes is important for the solution of a large variety of computer vision problems. While the majority of approaches has focused on static objects, dynamic objects, especially…

计算机视觉与模式识别 · 计算机科学 2022-05-18 Fangyin Wei , Rohan Chabra , Lingni Ma , Christoph Lassner , Michael Zollhöfer , Szymon Rusinkiewicz , Chris Sweeney , Richard Newcombe , Mira Slavcheva

Predicting human gaze in video is fundamental to advancing scene understanding and multimodal interaction. While traditional saliency maps provide spatial probability distributions and scanpaths offer ordered fixations, both abstractions…

计算机视觉与模式识别 · 计算机科学 2026-03-27 Jenna Kang , Colin Groth , Tong Wu , Finley Torrens , Patsorn Sangkloy , Gordon Wetzstein , Qi Sun