中文
相关论文

相关论文: EmbodiedOcc++: Boosting Embodied 3D Occupancy Pred…

200 篇论文

The lack of a large-scale 3D-text corpus has led recent works to distill open-vocabulary knowledge from vision-language models (VLMs). However, these methods typically rely on a single VLM to align the feature spaces of 3D models within a…

计算机视觉与模式识别 · 计算机科学 2025-03-31 Jinlong Li , Cristiano Saltori , Fabio Poiesi , Nicu Sebe

In this paper, we propose a geometric neural network with edge-aware refinement (GeoNet++) to jointly predict both depth and surface normal maps from a single image. Building on top of two-stream CNNs, GeoNet++ captures the geometric…

计算机视觉与模式识别 · 计算机科学 2020-12-15 Xiaojuan Qi , Zhengzhe Liu , Renjie Liao , Philip H. S. Torr , Raquel Urtasun , Jiaya Jia

To automatically localize a target object in an image is crucial for many computer vision applications. To represent the 2D object, ellipse labels have recently been identified as a promising alternative to axis-aligned bounding boxes. This…

计算机视觉与模式识别 · 计算机科学 2023-08-03 Vincent Gaudillière , Leo Pauly , Arunkumar Rathinam , Albert Garcia Sanchez , Mohamed Adel Musallam , Djamila Aouada

Medical images like CT and MRI provide detailed information about the internal structure of the body, and identifying key anatomical structures from these images plays a crucial role in clinical workflows. Current methods treat it as a…

计算机视觉与模式识别 · 计算机科学 2023-06-27 Xiaoyu Bai , Yong Xia

Weakly-supervised 3D occupancy perception is crucial for vision-based autonomous driving in outdoor environments. Previous methods based on NeRF often face a challenge in balancing the number of samples used. Too many samples can decrease…

计算机视觉与模式识别 · 计算机科学 2025-12-23 Qianpu Sun , Changyong Shu , Sifan Zhou , Runxi Cheng , Yongxian Wei , Zichen Yu , Dawei Yang , Sirui Han , Yuan Chun

Estimating the layout of a room from a single-shot panoramic image is important in virtual/augmented reality and furniture layout simulation. This involves identifying three-dimensional (3D) geometry, such as the location of corners and…

计算机视觉与模式识别 · 计算机科学 2023-04-26 Mizuki Tabata , Kana Kurata , Junichiro Tamamatsu

Vision-based occupancy prediction, also known as 3D Semantic Scene Completion (SSC), presents a significant challenge in computer vision. Previous methods, confined to onboard processing, struggle with simultaneous geometric and semantic…

计算机视觉与模式识别 · 计算机科学 2025-05-20 Hao Shi , Song Wang , Jiaming Zhang , Xiaoting Yin , Guangming Wang , Jianke Zhu , Kailun Yang , Kaiwei Wang

3D instance segmentation, with a variety of applications in robotics and augmented reality, is in large demands these days. Unlike 2D images that are projective observations of the environment, 3D models provide metric reconstruction of the…

计算机视觉与模式识别 · 计算机科学 2020-04-29 Lei Han , Tian Zheng , Lan Xu , Lu Fang

Recent advances in multimodal large language models (MLLMs) have opened new opportunities for embodied intelligence, enabling multimodal understanding, reasoning, and interaction, as well as continuous spatial decision-making. Nevertheless,…

Inferring the 3D structure of a scene from a single image is an ill-posed and challenging problem in the field of vision-centric autonomous driving. Existing methods usually employ neural radiance fields to produce voxelized 3D occupancy,…

计算机视觉与模式识别 · 计算机科学 2025-01-13 Yi Feng , Yu Han , Xijing Zhang , Tanghui Li , Yanting Zhang , Rui Fan

This paper presents a novel framework combining group equivariant convolutional neural networks (G-CNNs) with equivariant-aware structured pruning to produce compact, transformation-invariant models for resource-constrained environments.…

计算机视觉与模式识别 · 计算机科学 2025-11-24 Mohammed Alnemari

Spatial intelligence, encompassing 3D reconstruction, perception, and reasoning, is fundamental to applications such as robotics, aerial imaging, and extended reality. A key enabler is the real-time, accurate estimation of core 3D…

计算机视觉与模式识别 · 计算机科学 2025-07-11 Wenyan Cong , Yiqing Liang , Yancheng Zhang , Ziyi Yang , Yan Wang , Boris Ivanovic , Marco Pavone , Chen Chen , Zhangyang Wang , Zhiwen Fan

Recent advances in 3D Gaussian Splatting (3DGS) have enabled real-time, photorealistic scene reconstruction. However, conventional 3DGS frameworks typically rely on sparse point clouds derived from Structure-from-Motion (SfM), which…

图形学 · 计算机科学 2026-03-25 Yan Fang , Jianfei Ge , Jiangjian Xiao

Recognition of floor plans has been a challenging and popular task. Despite that many recent approaches have been proposed for this task, they typically fail to make the room-level unified prediction. Specifically, multiple semantic…

计算机视觉与模式识别 · 计算机科学 2022-11-01 Zhangyu Wang , Ningyuan Sun

3D Gaussian splatting (3DGS) has recently demonstrated promising advancements in RGB-D online dense mapping. Nevertheless, existing methods excessively rely on per-pixel depth cues to perform map densification, which leads to significant…

计算机视觉与模式识别 · 计算机科学 2024-09-02 Meng Wang , Junyi Wang , Changqun Xia , Chen Wang , Yue Qi

Understanding dynamic 3D environments in a spatially continuous and temporally consistent manner is fundamental for robotics and autonomous driving. While recent advances in occupancy prediction provide a unified representation of scene…

计算机视觉与模式识别 · 计算机科学 2026-03-10 Yongzhi Lin , Kai Luo , Yuanfan Zheng , Hao Shi , Mengfei Duan , Yang Liu , Kailun Yang

Single-image 3D reconstruction with large reconstruction models (LRMs) has advanced rapidly, yet reconstructions often exhibit geometric inconsistencies and misaligned details that limit fidelity. We introduce GeoFusionLRM, a geometry-aware…

计算机视觉与模式识别 · 计算机科学 2026-02-17 Ahmet Burak Yildirim , Tuna Saygin , Duygu Ceylan , Aysegul Dundar

Language-guided grasping has emerged as a promising paradigm for enabling robots to identify and manipulate target objects through natural language instructions, yet it remains highly challenging in cluttered or occluded scenes. Existing…

机器人学 · 计算机科学 2026-02-05 Rui Tang , Guankun Wang , Long Bai , Huxin Gao , Jiewen Lai , Chi Kit Ng , Jiazheng Wang , Fan Zhang , Hongliang Ren

Recent diffusion models have demonstrated remarkable performance in both 3D scene generation and perception tasks. Nevertheless, existing methods typically separate these two processes, acting as a data augmenter to generate synthetic data…

计算机视觉与模式识别 · 计算机科学 2025-08-25 Bohan Li , Xin Jin , Jianan Wang , Yukai Shi , Yasheng Sun , Xiaofeng Wang , Zhuang Ma , Baao Xie , Chao Ma , Xiaokang Yang , Wenjun Zeng

Tracking body and hand motions in the 3D space is essential for social and self-presence in augmented and virtual environments. Unlike the popular 3D pose estimation setting, the problem is often formulated as inside-out tracking based on…

计算机视觉与模式识别 · 计算机科学 2020-12-08 Mathias Parger , Chengcheng Tang , Yuanlu Xu , Christopher Twigg , Lingling Tao , Yijing Li , Robert Wang , Markus Steinberger