English
Related papers

Related papers: DINO-Foresight: Looking into the Future with DINO

200 papers

Foundation models (FMs) have revolutionized computer vision, enabling effective learning across different domains. However, their performance under domain shift is yet underexplored. This paper investigates the zero-shot domain adaptation…

Computer Vision and Pattern Recognition · Computer Science 2024-07-11 Ugur Ali Kaplan , Margret Keuper , Anna Khoreva , Dan Zhang , Yumeng Li

In this paper we present Mask DINO, a unified object detection and segmentation framework. Mask DINO extends DINO (DETR with Improved Denoising Anchor Boxes) by adding a mask prediction branch which supports all image segmentation tasks…

Computer Vision and Pattern Recognition · Computer Science 2022-12-13 Feng Li , Hao Zhang , Huaizhe xu , Shilong Liu , Lei Zhang , Lionel M. Ni , Heung-Yeung Shum

We present DINO Patch Visual Odometry (DINO-VO), an end-to-end monocular visual odometry system with strong scene generalization. Current Visual Odometry (VO) systems often rely on heuristic feature extraction strategies, which can degrade…

Computer Vision and Pattern Recognition · Computer Science 2026-04-07 Qi Chen , Guanghao Li , Sijia Hu , Xin Gao , Junpeng Ma , Xiangyang Xue , Jian Pu

Deep learning underlies most modern approaches and tools in computer vision, including biomedical imaging. However, for interactive semantic segmentation (often called pixel classification in this context) and interactive object-level…

Computer Vision and Pattern Recognition · Computer Science 2026-03-23 Carolin Teuber , Anwai Archit , Tobias Boothe , Peter Ditte , Jochen Rink , Constantin Pape

Surround-view depth estimation is a crucial task aims to acquire the depth maps of the surrounding views. It has many applications in real world scenarios such as autonomous driving, AR/VR and 3D reconstruction, etc. However, given that…

Computer Vision and Pattern Recognition · Computer Science 2024-05-28 Yifan Mao , Ming Li , Jian Liu , Jiayang Liu , Zihan Qin , Chunxi Chu , Jialei Xu , Wenbo Zhao , Junjun Jiang , Xianming Liu

Humans and animals have a rich and flexible understanding of the physical world, which enables them to infer the underlying dynamical trajectories of objects and events, plausible future states, and use that to plan and anticipate the…

Artificial Intelligence · Computer Science 2023-10-26 Aran Nayebi , Rishi Rajalingham , Mehrdad Jazayeri , Guangyu Robert Yang

Understanding the dynamics of generic 3D scenes is fundamentally challenging in computer vision, essential in enhancing applications related to scene reconstruction, motion tracking, and avatar creation. In this work, we address the task as…

Computer Vision and Pattern Recognition · Computer Science 2024-06-07 Yan Zhang , Sergey Prokudin , Marko Mihajlovic , Qianli Ma , Siyu Tang

What does it mean for a visual system to truly understand affordance? We argue that this understanding hinges on two complementary capacities: geometric perception, which identifies the structural parts of objects that enable interaction,…

Computer Vision and Pattern Recognition · Computer Science 2026-03-17 Qing Zhang , Xuesong Li , Jing Zhang

There is an inherent need for autonomous cars, drones, and other robots to have a notion of how their environment behaves and to anticipate changes in the near future. In this work, we focus on anticipating future appearance given the…

Computer Vision and Pattern Recognition · Computer Science 2017-07-25 Vedran Vukotić , Silvia-Laura Pintea , Christian Raymond , Guillaume Gravier , Jan Van Gemert

Open-Vocabulary Segmentation (OVS) aims to segment image regions beyond predefined category sets by leveraging semantic descriptions. While CLIP based approaches excel in semantic generalization, they frequently lack the fine-grained…

Computer Vision and Pattern Recognition · Computer Science 2026-04-10 Haoxi Zeng , Qiankun Liu , Yi Bin , Haiyue Zhang , Yujuan Ding , Guoqing Wang , Deqiang Ouyang , Heng Tao Shen

Visual Foundation Models (VFMs), such as DINO and CLIP, excel in semantic understanding of images but exhibit limited spatial reasoning capabilities, which limits their applicability to embodied systems. As a result, recent work…

Computer Vision and Pattern Recognition · Computer Science 2026-01-21 Turhan Can Kargin , Wojciech Jasiński , Adam Pardyl , Bartosz Zieliński , Marcin Przewięźlikowski

Forecasting from partial observations is central to world modeling. Many recent methods represent the world through images, and reduce forecasting to stochastic video generation. Although such methods excel at realism and visual fidelity,…

Computer Vision and Pattern Recognition · Computer Science 2025-12-15 Gabrijel Boduljak , Yushi Lan , Christian Rupprecht , Andrea Vedaldi

Face recognition systems are increasingly used in biometric security for convenience and effectiveness. However, they remain vulnerable to spoofing attacks, where attackers use photos, videos, or masks to impersonate legitimate users. This…

Computer Vision and Pattern Recognition · Computer Science 2024-10-23 Arman Keresh , Pakizar Shamoi

Although visual foundation models like DINOv2 provide state-of-the-art performance as feature extractors, their complex, high-dimensional representations create substantial hurdles for interpretability. This work proposes DINO-QPM, which…

Computer Vision and Pattern Recognition · Computer Science 2026-04-09 Robert Zimmermann , Thomas Norrenbrock , Bodo Rosenhahn

Although large-scale visual foundation models (VFMs) achieve remarkable performance in semantic understanding, they still underperform in instance-aware dense prediction tasks. They exhibit different biases in representation: for instance,…

Computer Vision and Pattern Recognition · Computer Science 2026-05-19 Yachan Guo , JoseLuis Gomez Zurita , Danna Xue , Yi Xiao , AntonioManuel Lopez Pena

Predicting the future trajectory of agents from visual observations is an important problem for realization of safe and effective navigation of autonomous systems in dynamic environments. This paper focuses on two important aspects of…

Computer Vision and Pattern Recognition · Computer Science 2020-07-24 Srikanth Malla , Isht Dwivedi , Behzad Dariush , Chiho Choi

Dense semantic forecasting anticipates future events in video by inferring pixel-level semantics of an unobserved future image. We present a novel approach that is applicable to various single-frame architectures and tasks. Our approach…

Computer Vision and Pattern Recognition · Computer Science 2022-01-06 Josip Šarić , Sacha Vražić , Siniša Šegvić

Anticipating future events is an important prerequisite towards intelligent behavior. Video forecasting has been studied as a proxy task towards this goal. Recent work has shown that to predict semantic segmentation of future frames,…

Computer Vision and Pattern Recognition · Computer Science 2018-10-04 Pauline Luc , Camille Couprie , Yann LeCun , Jakob Verbeek

There is substantial interest in developing artificial intelligence systems to support radiologists across tasks ranging from segmentation to report generation. Existing computed tomography (CT) foundation models have largely focused on…

Computer Vision and Pattern Recognition · Computer Science 2026-04-07 Rubén Moreno-Aguado , Alba Magallón , Victor Moreno , Yingying Fang , Guang Yang

Event cameras offer unique advantages for vision tasks in challenging environments, yet processing asynchronous event streams remains an open challenge. While existing methods rely on specialized architectures or resource-intensive…

Computer Vision and Pattern Recognition · Computer Science 2025-11-11 Ruihao Xia , Junhong Cai , Luziwei Leng , Liuyi Wang , Chengju Liu , Ran Cheng , Yang Tang , Pan Zhou