中文
相关论文

相关论文: FLEG: Feed-Forward Language Embedded Gaussian Spla…

200 篇论文

The semantically interactive radiance field has always been an appealing task for its potential to facilitate user-friendly and automated real-world 3D scene understanding applications. However, it is a challenging task to achieve high…

计算机视觉与模式识别 · 计算机科学 2024-12-13 Yuzhou Ji , He Zhu , Junshu Tang , Wuyi Liu , Zhizhong Zhang , Xin Tan , Yuan Xie

Feed-forward 3D reconstruction from sparse, low-resolution (LR) images is a crucial capability for real-world applications, such as autonomous driving and embodied AI. However, existing methods often fail to recover fine texture details.…

计算机视觉与模式识别 · 计算机科学 2025-11-18 Xinyuan Hu , Changyue Shi , Chuxiao Yang , Minghao Chen , Jiajun Ding , Tao Wei , Chen Wei , Zhou Yu , Min Tan

Reconstructing and segmenting scenes from unconstrained photo collections obtained from the Internet is a novel but challenging task. Unconstrained photo collections are easier to get than well-captured photo collections. These…

计算机视觉与模式识别 · 计算机科学 2025-07-11 Yongtang Bao , Chengjie Tang , Yuze Wang , Haojie Li

Injecting semantics into 3D Gaussian Splatting (3DGS) has recently garnered significant attention. While current approaches typically distill 3D semantic features from 2D foundational models (e.g., CLIP and SAM) to facilitate novel view…

计算机视觉与模式识别 · 计算机科学 2025-05-20 Wenbo Zhang , Lu Zhang , Ping Hu , Liqian Ma , Yunzhi Zhuge , Huchuan Lu

We propose DrivingForward, a feed-forward Gaussian Splatting model that reconstructs driving scenes from flexible surround-view input. Driving scene images from vehicle-mounted cameras are typically sparse, with limited overlap, and the…

计算机视觉与模式识别 · 计算机科学 2024-12-24 Qijian Tian , Xin Tan , Yuan Xie , Lizhuang Ma

High-fidelity 3D reconstruction is critical for aerial inspection tasks such as infrastructure monitoring, structural assessment, and environmental surveying. While traditional photogrammetry techniques enable geometric modeling, they lack…

图形学 · 计算机科学 2025-05-26 Mahmoud Chick Zaouali , Todd Charter , Homayoun Najjaran

Modeling open-vocabulary language fields in 3D is essential for intuitive human-AI interaction and querying within physical environments. State-of-the-art approaches, such as LangSplat, leverage 3D Gaussian Splatting to efficiently…

计算机视觉与模式识别 · 计算机科学 2025-10-28 Pranav Saxena

Reconstructing and understanding 3D scenes from unposed sparse views in a feed-forward manner remains as a challenging task in 3D computer vision. Recent approaches use per-pixel 3D Gaussian Splatting for reconstruction, followed by a…

Feed-forward 3D reconstruction offers substantial runtime advantages over per-scene optimization, which remains slow at inference and often fragile under sparse views. However, existing feed-forward methods still have potential for further…

计算机视觉与模式识别 · 计算机科学 2026-02-27 Tianyu Chen , Wei Xiang , Kang Han , Yu Lu , Di Wu , Gaowen Liu , Ramana Rao Kompella

We consider the problem of novel view synthesis from unposed images in a single feed-forward. Our framework capitalizes on fast speed, scalability, and high-quality 3D reconstruction and view synthesis capabilities of 3DGS, where we further…

计算机视觉与模式识别 · 计算机科学 2025-07-25 Sunghwan Hong , Jaewoo Jung , Heeseong Shin , Jisang Han , Jiaolong Yang , Chong Luo , Seungryong Kim

We present TokenSplat, a feed-forward framework for joint 3D Gaussian reconstruction and camera pose estimation from unposed multi-view images. At its core, TokenSplat introduces a Token-aligned Gaussian Prediction module that aligns…

计算机视觉与模式识别 · 计算机科学 2026-03-03 Yihui Li , Chengxin Lv , Zichen Tang , Hongyu Yang , Di Huang

Constructing a 3D scene capable of accommodating open-ended language queries, is a pivotal pursuit, particularly within the domain of robotics. Such technology facilitates robots in executing object manipulations based on human language…

The emergence of neural representations has revolutionized our means for digitally viewing a wide range of 3D scenes, enabling the synthesis of photorealistic images rendered from novel views. Recently, several techniques have been proposed…

计算机视觉与模式识别 · 计算机科学 2025-02-14 Gal Fiebelman , Tamir Cohen , Ayellet Morgenstern , Peter Hedman , Hadar Averbuch-Elor

3D Gaussian Splatting (3DGS) has demonstrated impressive novel view synthesis performance. While conventional methods require per-scene optimization, more recently several feed-forward methods have been proposed to generate pixel-aligned…

计算机视觉与模式识别 · 计算机科学 2025-03-21 Shengjun Zhang , Xin Fei , Fangfu Liu , Haixu Song , Yueqi Duan

3D semantic field learning is crucial for applications like autonomous navigation, AR/VR, and robotics, where accurate comprehension of 3D scenes from limited viewpoints is essential. Existing methods struggle under sparse view conditions,…

计算机视觉与模式识别 · 计算机科学 2025-08-19 Kangjie Chen , BingQuan Dai , Minghan Qin , Dongbin Zhang , Peihao Li , Yingshuang Zou , Haoqian Wang

While existing feed-forward Gaussian splatting models offer computational efficiency and can generalize to sparse view settings, their performance is fundamentally constrained by relying on a single forward pass for inference. We propose…

计算机视觉与模式识别 · 计算机科学 2026-03-13 Haofei Xu , Daniel Barath , Andreas Geiger , Marc Pollefeys

We present FPGS, a feed-forward photorealistic style transfer method of large-scale radiance fields represented by Gaussian Splatting. FPGS, stylizes large-scale 3D scenes with arbitrary, multiple style reference images without additional…

图形学 · 计算机科学 2025-03-14 GeonU Kim , Kim Youwang , Lee Hyoseok , Tae-Hyun Oh

We present Splat-SAP, a feed-forward approach to render novel views of human-centered scenes from binocular cameras with large sparsity. Gaussian Splatting has shown its promising potential in rendering tasks, but it typically necessitates…

计算机视觉与模式识别 · 计算机科学 2025-12-01 Boyao Zhou , Shunyuan Zheng , Zhanfeng Liao , Zihan Ma , Hanzhang Tu , Boning Liu , Yebin Liu

3D open-vocabulary scene understanding, which accurately perceives complex semantic properties of objects in space, has gained significant attention in recent years. In this paper, we propose GAGS, a framework that distills 2D CLIP features…

计算机视觉与模式识别 · 计算机科学 2025-03-11 Yuning Peng , Haiping Wang , Yuan Liu , Chenglu Wen , Zhen Dong , Bisheng Yang

3D Gaussian Splatting (3DGS) has emerged as a transformative method in the field of real-time novel synthesis. Based on 3DGS, recent advancements cope with large-scale scenes via spatial-based partition strategy to reduce video memory and…

计算机视觉与模式识别 · 计算机科学 2025-01-06 Tengfei Wang , Xin Wang , Yongmao Hou , Yiwei Xu , Wendi Zhang , Zongqian Zhan