中文
相关论文

相关论文: Generating Visual Spatial Description via Holistic…

200 篇论文

Advancements in 3D rendering like Gaussian Splatting (GS) allow novel view synthesis and real-time rendering in virtual reality (VR). However, GS-created 3D environments are often difficult to edit. For scene enhancement or to incorporate…

计算机视觉与模式识别 · 计算机科学 2024-09-25 Hannah Schieber , Jacob Young , Tobias Langlotz , Stefanie Zollmann , Daniel Roth

Scene representations using 3D Gaussian primitives have produced excellent results in modeling the appearance of static and dynamic 3D scenes. Many graphics applications, however, demand the ability to manipulate both the appearance and the…

计算机视觉与模式识别 · 计算机科学 2024-04-02 Ri-Zhao Qiu , Ge Yang , Weijia Zeng , Xiaolong Wang

Understanding 3D scenes goes beyond simply recognizing objects; it requires reasoning about the spatial and semantic relationships between them. Current 3D scene-language models often struggle with this relational understanding,…

计算机视觉与模式识别 · 计算机科学 2025-12-09 Jintang Xue , Ganning Zhao , Jie-En Yao , Hong-En Chen , Yue Hu , Meida Chen , Suya You , C. -C. Jay Kuo

Deep generative models allow for photorealistic image synthesis at high resolutions. But for many applications, this is not enough: content creation also needs to be controllable. While several recent works investigate how to disentangle…

计算机视觉与模式识别 · 计算机科学 2021-04-30 Michael Niemeyer , Andreas Geiger

Text-based Visual Question Answering~(TextVQA) aims to produce correct answers for given questions about the images with multiple scene texts. In most cases, the texts naturally attach to the surface of the objects. Therefore, spatial…

计算机视觉与模式识别 · 计算机科学 2023-06-16 Hao Li , Jinfa Huang , Peng Jin , Guoli Song , Qi Wu , Jie Chen

Open-vocabulary 3D Scene Graph (3DSG) can enhance various downstream tasks in robotics by leveraging structured semantic representations, yet current 3DSG construction methods suffer from semantic inconsistencies caused by noisy cross-image…

计算机视觉与模式识别 · 计算机科学 2026-03-17 Yue Chang , Rufeng Chen , Zhaofan Zhang , Yi Chen , Yifan Tian , Sihong Xie

3D Gaussian Splatting (3DGS) has recently emerged as an innovative and efficient 3D representation technique. While its potential for extended reality (XR) applications is frequently highlighted, its practical effectiveness remains…

计算机视觉与模式识别 · 计算机科学 2025-01-17 Shi Qiu , Binzhu Xie , Qixuan Liu , Pheng-Ann Heng

Existing 3D reconstruction methods utilize guidances such as 2D images, 3D point clouds, shape contours and single semantics to recover the 3D surface, which limits the creative exploration of 3D modeling. In this paper, we propose a novel…

计算机视觉与模式识别 · 计算机科学 2025-08-19 Liangchen Li , Caoliwen Wang , Yuqi Zhou , Bailin Deng , Juyong Zhang

3D Gaussian Splatting (3DGS) has recently gained popularity for efficient scene rendering by representing scenes as explicit sets of anisotropic 3D Gaussians. However, most existing work focuses primarily on modeling external surfaces. In…

图像与视频处理 · 电气工程与系统科学 2026-01-12 Shuxin Liang , Yihan Xiao , Wenlu Tang

3D Gaussian Splatting SLAM has emerged as a widely used technique for high-fidelity mapping in spatial intelligence. However, existing methods often rely on a single representation scheme, which limits their performance in large-scale…

计算机视觉与模式识别 · 计算机科学 2025-10-28 Wenkai Zhu , Xu Li , Qimin Xu , Benwu Wang , Kun Wei , Yiming Peng , Zihang Wang

Generating 4D scenes from a single-view video is inherently ill-posed: a single viewpoint lacks the information needed to recover a complete, dynamic scene with full coverage. Existing methods are typically limited to monocular videos,…

计算机视觉与模式识别 · 计算机科学 2026-05-26 Tingxi Chen , Ke Hao , Yabo Chen , Zhengxue Cheng , Rong Xie , Li Song , Haibin Huang , Chi Zhang , Xuelong Li

Vision-language models (VLMs) have advanced multimodal reasoning but still face challenges in spatial reasoning for 3D scenes and complex object configurations. To address this, we introduce SpatialViLT, an enhanced VLM that integrates…

计算机视觉与模式识别 · 计算机科学 2025-10-07 Chashi Mahiul Islam , Oteo Mamo , Samuel Jacob Chacko , Xiuwen Liu , Weikuan Yu

Camera-based 3D Semantic Scene Completion (SSC) is a critical task for autonomous driving and robotic scene understanding. It aims to infer a complete 3D volumetric representation of both semantics and geometry from a single image. Existing…

计算机视觉与模式识别 · 计算机科学 2026-04-07 Zaidao Han , Risa Higashita , Jiang Liu

Due to the complex and highly dynamic motions in the real world, synthesizing dynamic videos from multi-view inputs for arbitrary viewpoints is challenging. Previous works based on neural radiance field or 3D Gaussian splatting are limited…

计算机视觉与模式识别 · 计算机科学 2025-07-04 Jiahao Wu , Rui Peng , Jianbo Jiao , Jiayu Yang , Luyang Tang , Kaiqiang Xiong , Jie Liang , Jinbo Yan , Runling Liu , Ronggang Wang

Current approaches for open-vocabulary scene graph generation (OVSGG) use vision-language models such as CLIP and follow a standard zero-shot pipeline -- computing similarity between the query image and the text embeddings for each category…

计算机视觉与模式识别 · 计算机科学 2024-10-22 Guikun Chen , Jin Li , Wenguan Wang

Most deep learning approaches to comprehensive semantic modeling of 3D indoor spaces require costly dense annotations in the 3D domain. In this work, we explore a central 3D scene modeling task, namely, semantic scene reconstruction without…

计算机视觉与模式识别 · 计算机科学 2024-06-06 Junwen Huang , Alexey Artemov , Yujin Chen , Shuaifeng Zhi , Kai Xu , Matthias Nießner

Dynamic scene rendering opens new avenues in autonomous driving by enabling closed-loop simulations with photorealistic data, which is crucial for validating end-to-end algorithms. However, the complex and highly dynamic nature of traffic…

计算机视觉与模式识别 · 计算机科学 2025-03-11 Rui Song , Chenwei Liang , Yan Xia , Walter Zimmer , Hu Cao , Holger Caesar , Andreas Festag , Alois Knoll

Functional 3D scene graphs offer a versatile and flexible representation for 3D scene understanding and robotic manipulation, defined by object nodes, interactive elements, and functional relationship edges. However, their potential remains…

Reconstructing coherent 3D geometry and appearance from unposed multi-view images is a fundamental yet challenging problem in computer vision. Most existing visual geometry foundation models predict explicit geometry by regressing…

计算机视觉与模式识别 · 计算机科学 2026-05-22 Yuqi Wu , Tianyu Hu , Wenzhao Zheng , Yuanhui Huang , Haowen Sun , Jie Zhou , Jiwen Lu

Traffic scene understanding is essential for enabling autonomous vehicles to accurately perceive and interpret their environment, thereby ensuring safe navigation. This paper presents a novel framework that transforms a single frontal-view…

计算机视觉与模式识别 · 计算机科学 2026-01-22 Danial Sadrian Zadeh , Otman A. Basir , Behzad Moshiri