中文
相关论文

相关论文: Light Cones For Vision: Simple Causal Priors For V…

200 篇论文

The de facto approach in video object-centric learning maintains temporal consistency through learned dynamics modules that predict future object representations, called slots. We demonstrate that these predictors function as expensive…

计算机视觉与模式识别 · 计算机科学 2026-05-12 Zhiyuan Li , Rongzhen Zhao , Wenyan Yang , Wenshuai Zhao , Pekka Marttinen , Joni Pajarinen

We show that four-dimensional Lorentzian metrics admitting a global spacelike Lie group of isometries, $G_{1}={\mathbb R}$, which obey the Einstein equations for vacuum and certain types of matter, cannot contain apparent horizons. The…

广义相对论与量子宇宙学 · 物理学 2007-05-23 Sergio M. C. V. Goncalves

Weakly supervised localization aims at finding target object regions using only image-level supervision. However, localization maps extracted from classification networks are often not accurate due to the lack of fine pixel-level…

计算机视觉与模式识别 · 计算机科学 2020-08-13 Xiaolin Zhang , Yunchao Wei , Yi Yang

Shortcut learning, where machine learning models exploit spurious correlations in data instead of capturing meaningful features, poses a significant challenge to building robust and generalizable models. This phenomenon is prevalent across…

机器学习 · 计算机科学 2025-09-03 Pirzada Suhail , Vrinda Goel , Amit Sethi

Hyperbolic space naturally encodes hierarchical structures such as phylogenies (binary trees), where inward-bending geodesics reflect paths through least common ancestors, and the exponential growth of neighborhoods mirrors the…

3D object detection from a single image without LiDAR is a challenging task due to the lack of accurate depth information. Conventional 2D convolutions are unsuitable for this task because they fail to capture local object and its scale…

计算机视觉与模式识别 · 计算机科学 2019-12-16 Mingyu Ding , Yuqi Huo , Hongwei Yi , Zhe Wang , Jianping Shi , Zhiwu Lu , Ping Luo

Retinal image of surrounding objects varies tremendously due to the changes in position, size, pose, illumination condition, background context, occlusion, noise, and nonrigid deformations. But despite these huge variations, our visual…

计算机视觉与模式识别 · 计算机科学 2017-02-14 Saeed Reza Kheradpisheh , Mohammad Ganjtabesh , Timothée Masquelier

Deep neural networks trained with different architectures, objectives, and datasets have been reported to converge on similar visual representations. However, what remains unknown is which visual properties models actually converge on and…

计算机视觉与模式识别 · 计算机科学 2026-05-14 Florian P. Mahner , Johannes Roth , Ka Chun Lam , Michael F. Bonner , Francisco Pereira , Martin N. Hebart

Humans can easily infer the underlying 3D geometry and texture of an object only from a single 2D image. Current computer vision methods can do this, too, but suffer from view generalization problems - the models inferred tend to make poor…

计算机视觉与模式识别 · 计算机科学 2021-06-14 Anand Bhattad , Aysegul Dundar , Guilin Liu , Andrew Tao , Bryan Catanzaro

We consider problems in model selection caused by the geometry of models close to their points of intersection. In some cases---including common classes of causal or graphical models, as well as time series models---distinct models may…

统计理论 · 数学 2022-12-20 Robin J. Evans

The seemingly infinite diversity of the natural world arises from a relatively small set of coherent rules, such as the laws of physics or chemistry. We conjecture that these rules give rise to regularities that can be discovered through…

Building successful recommender systems requires uncovering the underlying dimensions that describe the properties of items as well as users' preferences toward them. In domains like clothing recommendation, explaining users' preferences…

信息检索 · 计算机科学 2016-04-21 Ruining He , Chunbin Lin , Jianguo Wang , Julian McAuley

Vision Language Models (VLMs) excel at identifying and describing objects but often fail at spatial reasoning. We study why VLMs, such as LLaVA, underutilize spatial cues despite having positional encodings and spatially rich vision encoder…

计算机视觉与模式识别 · 计算机科学 2025-10-02 Jianing Qi , Jiawei Liu , Hao Tang , Zhigang Zhu

This paper proves that visual object recognition systems using only 2D Euclidean similarity measurements to compare object views against previously seen views can achieve the same recognition performance as observers having access to all…

计算机视觉与模式识别 · 计算机科学 2007-12-04 Thomas M. Breuel

Our brain can almost effortlessly decompose visual data streams into background and salient objects. Moreover, it can anticipate object motion and interactions, which are crucial abilities for conceptual planning and reasoning. Recent…

计算机视觉与模式识别 · 计算机科学 2023-02-08 Manuel Traub , Sebastian Otte , Tobias Menge , Matthias Karlbauer , Jannik Thümmel , Martin V. Butz

We prove an exponential separation in sample complexity between Euclidean and hyperbolic representations for learning on hierarchical data under standard Lipschitz regularization. For depth-$R$ hierarchies with branching factor $m$, we…

机器学习 · 统计学 2026-01-29 Divit Rawal , Sriram Vishwanath

The Lorentzian metric structure used in any field theory allows one to implement the relativistic notion of causality and to define a notion of time dimension. This article investigates the possibility that at the microscopic level the…

高能物理 - 理论 · 物理学 2013-04-11 Shinji Mukohyama , Jean-Philippe Uzan

Inferring geometrically consistent dense 3D scenes across a tuple of temporally consecutive images remains challenging for self-supervised monocular depth prediction pipelines. This paper explores how the increasingly popular transformer…

计算机视觉与模式识别 · 计算机科学 2021-10-18 Patrick Ruhkamp , Daoyi Gao , Hanzhi Chen , Nassir Navab , Benjamin Busam

The geodesic-light-cone (GLC) coordinates are a useful tool to analyse light propagation and observations in cosmological models. In this article, we propose a detailed, pedagogical, and rigorous introduction to this coordinate system,…

广义相对论与量子宇宙学 · 物理学 2016-06-09 Pierre Fleury , Fabien Nugier , Giuseppe Fanizza

Category-level object pose estimation aims to determine the pose and size of novel objects in specific categories. Existing correspondence-based approaches typically adopt point-based representations to establish the correspondences between…

计算机视觉与模式识别 · 计算机科学 2025-03-20 Huan Ren , Wenfei Yang , Xiang Liu , Shifeng Zhang , Tianzhu Zhang