English
Related papers

Related papers: PLAF: Pixel-wise Language-Aligned Feature Extracti…

200 papers

Diffusion models have recently received increasing research attention for their remarkable transfer abilities in semantic segmentation tasks. However, generating fine-grained segmentation masks with diffusion models often requires…

Computer Vision and Pattern Recognition · Computer Science 2024-01-23 Koichi Namekata , Amirmojtaba Sabour , Sanja Fidler , Seung Wook Kim

The semantically interactive radiance field has long been a promising backbone for 3D real-world applications, such as embodied AI to achieve scene understanding and manipulation. However, multi-granularity interaction remains a challenging…

Computer Vision and Pattern Recognition · Computer Science 2025-04-14 Xin Tan , Yuzhou Ji , He Zhu , Yuan Xie

Open-vocabulary 3D object detection for autonomous driving aims to detect novel objects beyond the predefined training label sets in point cloud scenes. Existing approaches achieve this by connecting traditional 3D object detectors with…

Computer Vision and Pattern Recognition · Computer Science 2025-03-11 Adrian Chow , Evelien Riddell , Yimu Wang , Sean Sedwards , Krzysztof Czarnecki

Indoor scene understanding remains a fundamental challenge in robotics, with direct implications for downstream tasks such as navigation and manipulation. Traditional approaches often rely on closed-set recognition or loop closure, limiting…

Robotics · Computer Science 2025-06-10 Hongming Chen , Yiyang Lin , Ziliang Li , Biyu Ye , Yuying Zhang , Ximin Lyu

Recent advances in vision-language models (VLMs) have garnered substantial attention in open-vocabulary semantic and part segmentation (OSPS). However, existing methods extract image-text alignment cues from cost volumes through a serial…

Computer Vision and Pattern Recognition · Computer Science 2026-03-19 Jianjian Yin , Tao Chen , Yi Chen , Gensheng Pei , Xiangbo Shu , Yazhou Yao , Fumin Shen

Emerging neural radiance fields (NeRF) are a promising scene representation for computer graphics, enabling high-quality 3D reconstruction and novel view synthesis from image observations. However, editing a scene represented by a NeRF is…

Computer Vision and Pattern Recognition · Computer Science 2022-10-18 Sosuke Kobayashi , Eiichi Matsumoto , Vincent Sitzmann

While Large language models (LLMs) have advanced natural language processing tasks, their growing computational and memory demands make deployment on resource-constrained devices like mobile phones increasingly challenging. In this paper,…

Machine Learning · Computer Science 2025-02-13 Yiping Wang , Hanxian Huang , Yifang Chen , Jishen Zhao , Simon Shaolei Du , Yuandong Tian

We study open-world 3D scene understanding, a family of tasks that require agents to reason about their 3D environment with an open-set vocabulary and out-of-domain visual inputs - a critical skill for robots to operate in the unstructured…

Computer Vision and Pattern Recognition · Computer Science 2022-12-07 Huy Ha , Shuran Song

The growing demands of artificial intelligence and immersive media require communication beyond bit-level accuracy to meaning awareness. Conventional optical systems that focused on syntactic precision suffer significant inefficiencies.…

We generalize lifting to semantic lifting by incorporating per-view masks that indicate relevant pixels for lifting tasks. These masks are determined by querying corresponding multiscale pixel-aligned feature maps, which are derived from…

Computer Vision and Pattern Recognition · Computer Science 2025-03-11 George Tang , Aditya Agarwal , Weiqiao Han , Trevor Darrell , Yutong Bai

This paper describes a fast and accurate semantic image segmentation approach that encodes not only the discriminative features from deep neural networks, but also the high-order context compatibility among adjacent objects as well as low…

Computer Vision and Pattern Recognition · Computer Science 2016-05-16 Falong Shen , Gang Zeng

Open-vocabulary object detection (OVD) has been studied with Vision-Language Models (VLMs) to detect novel objects beyond the pre-trained categories. Previous approaches improve the generalization ability to expand the knowledge of the…

Computer Vision and Pattern Recognition · Computer Science 2024-04-09 Jooyeon Kim , Eulrang Cho , Sehyung Kim , Hyunwoo J. Kim

Open-vocabulary scene understanding is crucial for robotic applications, enabling robots to comprehend complex 3D environmental contexts and supporting various downstream tasks such as navigation and manipulation. However, existing methods…

Geometrically accurate and semantically expressive map representations have proven invaluable for robot deployment and task planning in unknown environments. Nevertheless, real-time, open-vocabulary semantic understanding of large-scale…

Recent progress of deep image classification models has provided great potential to improve state-of-the-art performance in related computer vision tasks. However, the transition to semantic segmentation is hampered by strict memory…

Computer Vision and Pattern Recognition · Computer Science 2019-05-15 Ivan Krešo , Josip Krapac , Siniša Šegvić

Modeling 3D language fields with Gaussian Splatting for open-ended language queries has recently garnered increasing attention. However, recent 3DGS-based models leverage view-dependent 2D foundation models to refine 3D semantics but lack a…

Computer Vision and Pattern Recognition · Computer Science 2025-06-10 Chenlu Zhan , Yufei Zhang , Gaoang Wang , Hongwei Wang

In this paper we present an efficient method for aggregating binary feature descriptors to allow compact representation of 3D scene model in incremental structure-from-motion and SLAM applications. All feature descriptors linked with one 3D…

Computer Vision and Pattern Recognition · Computer Science 2018-10-01 Jacek Komorowski , Tomasz Trzcinski

In this work, we pioneer Semantic Flow, a neural semantic representation of dynamic scenes from monocular videos. In contrast to previous NeRF methods that reconstruct dynamic scenes from the colors and volume densities of individual…

Computer Vision and Pattern Recognition · Computer Science 2024-04-09 Fengrui Tian , Yueqi Duan , Angtian Wang , Jianfei Guo , Shaoyi Du

Open-vocabulary 3D scene understanding is pivotal for enhancing physical intelligence, as it enables embodied agents to interpret and interact dynamically within real-world environments. This paper introduces MPEC, a novel Masked…

Computer Vision and Pattern Recognition · Computer Science 2025-04-29 Yan Wang , Baoxiong Jia , Ziyu Zhu , Siyuan Huang

We propose SLARM, a feed-forward model that unifies dynamic scene reconstruction, semantic understanding, and real-time streaming inference. SLARM captures complex, non-uniform motion through higher-order motion modeling, trained solely on…

Computer Vision and Pattern Recognition · Computer Science 2026-03-27 Zhicheng Qiu , Jiarui Meng , Tong-an Luo , Yican Huang , Xuan Feng , Xuanfu Li , ZHan Xu