English
Related papers

Related papers: Geometry Meets Vision: Revisiting Pretrained Seman…

200 papers

Visual grounding is a task that aims to locate a target object according to a natural language expression. As a multi-modal task, feature interaction between textual and visual inputs is vital. However, previous solutions mainly handle each…

Computer Vision and Pattern Recognition · Computer Science 2022-06-23 Chonghan Chen , Qi Jiang , Chih-Hao Wang , Noel Chen , Haohan Wang , Xiang Li , Bhiksha Raj

Unsupervised learning for geometric perception (depth, optical flow, etc.) is of great interest to autonomous systems. Recent works on unsupervised learning have made considerable progress on perceiving geometry; however, they usually…

Computer Vision and Pattern Recognition · Computer Science 2019-04-08 Yue Meng , Yongxi Lu , Aman Raj , Samuel Sunarjo , Rui Guo , Tara Javidi , Gaurav Bansal , Dinesh Bharadia

Self-supervised pre-training strategies have recently shown impressive results for training general-purpose feature extraction backbones in computer vision. In combination with the Vision Transformer architecture, the DINO self-distillation…

Computer Vision and Pattern Recognition · Computer Science 2023-04-18 Alexander Koenig , Maximilian Schambach , Johannes Otterbach

During 3D reconstruction, it is often the case that people cannot scan each individual object from all views, resulting in missing geometry in the captured scan. This missing geometry can be fundamentally limiting for many applications,…

Computer Vision and Pattern Recognition · Computer Science 2020-03-13 Ji Hou , Angela Dai , Matthias Nießner

Grounding open-ended semantic instructions into physically executable local goals is a fundamental challenge in human-robot interaction. While existing navigation frameworks often regress deterministic waypoints, this rigid formulation…

Robotics · Computer Science 2026-05-20 Kaijie Yun , Yue Chen

Completing a corrupted image with correct structures and reasonable textures for a mixed scene remains an elusive challenge. Since the missing hole in a mixed scene of a corrupted image often contains various semantic information,…

Computer Vision and Pattern Recognition · Computer Science 2020-07-13 Liang Liao , Jing Xiao , Zheng Wang , Chia-Wen Lin , Shin'ichi Satoh

Interpretability methods for deep neural networks mainly focus on the sensitivity of the class score with respect to the original or perturbed input, usually measured using actual or modified gradients. Some methods also use a…

Computer Vision and Pattern Recognition · Computer Science 2022-06-20 Md Mahfuzur Rahman , Noah Lewis , Sergey Plis

General object composition (GOC) aims to seamlessly integrate a target object into a background scene with desired geometric properties, while simultaneously preserving its fine-grained appearance details. Recent approaches derive semantic…

Computer Vision and Pattern Recognition · Computer Science 2026-05-19 Jianman Lin , Haojie Li , Chunmei Qing , Zhijing Yang , Liang Lin , Tianshui Chen

We tackle the task of learning dynamic 3D semantic radiance fields given a single monocular video as input. Our learned semantic radiance field captures per-point semantics as well as color and geometric properties for a dynamic 3D scene,…

Computer Vision and Pattern Recognition · Computer Science 2024-05-30 Isaac Labe , Noam Issachar , Itai Lang , Sagie Benaim

Instruction-based image editing (IIE) models have recently demonstrated strong capability in modifying specific image regions according to natural language instructions, which implicitly requires identifying where an edit should be applied.…

Computer Vision and Pattern Recognition · Computer Science 2026-05-14 Jingxuan He , Xiyu Wang , Yunke Wang , Mengyu Zheng , Chang Xu

Space grounding refers to localizing a set of spatial references described in natural language instructions. Traditional methods often fail to account for complex reasoning -- such as distance, geometry, and inter-object relationships --…

Robotics · Computer Science 2025-11-20 Nayoung Oh , Dohyun Kim , Junhyeong Bang , Rohan Paul , Daehyung Park

Visual Language Models (VLMs) have increasingly become the main paradigm for understanding indoor scenes, but they still struggle with metric and spatial reasoning. Current approaches rely on end-to-end video understanding or large-scale…

Computer Vision and Pattern Recognition · Computer Science 2026-03-17 Fernando Ropero , Erkin Turkoz , Daniel Matos , Junqing Du , Antonio Ruiz , Yanfeng Zhang , Lu Liu , Mingwei Sun , Yongliang Wang

Despite recent advancements in text-to-image diffusion models facilitating various image editing techniques, complex text prompts often lead to an oversight of some requests due to a bottleneck in processing text information. To tackle this…

Computer Vision and Pattern Recognition · Computer Science 2024-03-21 Hangeol Chang , Jinho Chang , Jong Chul Ye

Spatial intelligence requires visual representations that capture both semantic objects and geometric structure in the physical world. To support this, two major pre-training schemes are now widely used as foundation backbones:…

Computer Vision and Pattern Recognition · Computer Science 2026-05-28 Haozhan Shen , Tiancheng Zhao , Kangjia Zhao , Jianwei Yin

Reproducible closed-loop evaluation remains a major bottleneck in Embodied AI such as visual navigation. A promising path forward is high-fidelity simulation that combines photorealistic sensor rendering with geometrically grounded…

Computer Vision and Pattern Recognition · Computer Science 2026-03-30 Xinhao Liu , Jiaqi Li , Youming Deng , Ruxin Chen , Yingjia Zhang , Yifei Ma , Li Guo , Yiming Li , Jing Zhang , Chen Feng

Empowered by large-scale training, vision-language models (VLMs) achieve strong image and video understanding, yet their ability to perform spatial reasoning in both static scenes and dynamic videos remains limited. Recent advances try to…

Computer Vision and Pattern Recognition · Computer Science 2026-03-30 Shihua Zhang , Qiuhong Shen , Shizun Wang , Tianbo Pan , Xinchao Wang

Vision-Language Models (VLMs) have become indispensable for multimodal reasoning, yet their representations often encode and amplify demographic biases, resulting in biased associations and misaligned predictions in downstream tasks. Such…

Computer Vision and Pattern Recognition · Computer Science 2026-04-03 Dachuan Zhao , Weiyue Li , Zhenda Shen , Yushu Qiu , Bowen Xu , Haoyu Chen , Yongchao Chen

Visual grounding is a common vision task that involves grounding descriptive sentences to the corresponding regions of an image. Most existing methods use independent image-text encoding and apply complex hand-crafted modules or…

Computer Vision and Pattern Recognition · Computer Science 2024-10-29 Ming Dai , Lingfeng Yang , Yihao Xu , Zhenhua Feng , Wankou Yang

Large Vision-Language Models (LVLMs) offer remarkable benefits for a variety of vision-language tasks. However, a challenge hindering their application in real-world scenarios, particularly regarding safety, robustness, and reliability, is…

Computer Vision and Pattern Recognition · Computer Science 2024-01-17 Jiaying Lu , Jinmeng Rao , Kezhen Chen , Xiaoyuan Guo , Yawen Zhang , Baochen Sun , Carl Yang , Jie Yang

We introduce Grounded SAM, which uses Grounding DINO as an open-set object detector to combine with the segment anything model (SAM). This integration enables the detection and segmentation of any regions based on arbitrary text inputs and…

Computer Vision and Pattern Recognition · Computer Science 2024-01-26 Tianhe Ren , Shilong Liu , Ailing Zeng , Jing Lin , Kunchang Li , He Cao , Jiayu Chen , Xinyu Huang , Yukang Chen , Feng Yan , Zhaoyang Zeng , Hao Zhang , Feng Li , Jie Yang , Hongyang Li , Qing Jiang , Lei Zhang
‹ Prev 1 4 5 6 7 8 10 Next ›