English
Related papers

Related papers: SeqAffordSplat: Scene-level Sequential Affordance …

200 papers

Modeling 3D language fields with Gaussian Splatting for open-ended language queries has recently garnered increasing attention. However, recent 3DGS-based models leverage view-dependent 2D foundation models to refine 3D semantics but lack a…

Computer Vision and Pattern Recognition · Computer Science 2025-06-10 Chenlu Zhan , Yufei Zhang , Gaoang Wang , Hongwei Wang

This work presents IAAO, a novel framework that builds an explicit 3D model for intelligent agents to gain understanding of articulated objects in their environment through interaction. Unlike prior methods that rely on task-specific…

Computer Vision and Pattern Recognition · Computer Science 2025-04-10 Can Zhang , Gim Hee Lee

Embodied agents operating in open environments must translate high-level instructions into grounded, executable behaviors, often requiring coordinated use of both hands. While recent foundation models offer strong semantic reasoning,…

Robotics · Computer Science 2025-12-11 Kwang Bin Lee , Jiho Kang , Sung-Hee Lee

Recently, Gaussian Splatting, a method that represents a 3D scene as a collection of Gaussian distributions, has gained significant attention in addressing the task of novel view synthesis. In this paper, we highlight a fundamental…

Computer Vision and Pattern Recognition · Computer Science 2024-10-31 Haoxuan Qu , Zhuoling Li , Hossein Rahmani , Yujun Cai , Jun Liu

3D affordance grounding aims to highlight the actionable regions on 3D objects, which is crucial for robotic manipulation. Previous research primarily focused on learning affordance knowledge from static cues such as language and images,…

Computer Vision and Pattern Recognition · Computer Science 2026-02-11 Hanqing Wang , Mingyu Liu , Xiaoyu Chen , Chengwei MA , Yiming Zhong , Wenti Yin , Yuhao Liu , Zhiqing Cui , Jiahao Yuan , Lu Dai , Zhiyuan Ma , Hui Xiong

The emergence of 3D Gaussian Splatting (3DGS) has greatly accelerated the rendering speed of novel view synthesis. Unlike neural implicit representations like Neural Radiance Fields (NeRF) that represent a 3D scene with position and…

Computer Vision and Pattern Recognition · Computer Science 2024-04-16 Tong Wu , Yu-Jie Yuan , Ling-Xiao Zhang , Jie Yang , Yan-Pei Cao , Ling-Qi Yan , Lin Gao

3D Gaussian Splatting (3DGS) has gained significant attention for its real-time, photo-realistic rendering in novel-view synthesis and 3D modeling. However, existing methods struggle with accurately modeling in-the-wild scenes affected by…

Computer Vision and Pattern Recognition · Computer Science 2025-12-05 Chuanyu Fu , Guanying Chen , Yuqi Zhang , Kunbin Yao , Yuan Xiong , Chuan Huang , Shuguang Cui , Yasuyuki Matsushita , Xiaochun Cao

This paper presents a multimodal framework that integrates touch signals (contact points and surface normals) into 3D Gaussian Splatting (3DGS). Our approach enhances scene reconstruction, particularly under challenging conditions like low…

Signal Processing · Electrical Eng. & Systems 2025-08-12 Yuchen Gao , Xiao Xu , Eckehard Steinbach , Daniel E. Lucani , Qi Zhang

In this work, we present Fed3DGS, a scalable 3D reconstruction framework based on 3D Gaussian splatting (3DGS) with federated learning. Existing city-scale reconstruction methods typically adopt a centralized approach, which gathers all…

Computer Vision and Pattern Recognition · Computer Science 2024-03-19 Teppei Suzuki

3D Gaussian Splatting (3DGS) has revolutionized fast novel view synthesis, yet its opacity-based formulation makes surface extraction fundamentally difficult. Unlike implicit methods built on Signed Distance Fields or occupancy, 3DGS lacks…

Computer Vision and Pattern Recognition · Computer Science 2026-04-09 Diego Gomez , Antoine Guédon , Nissim Maruani , Bingchen Gong , Maks Ovsjanikov

A major breakthrough in 3D reconstruction is the feedforward paradigm to generate pixel-wise 3D points or Gaussian primitives from sparse, unposed images. To further incorporate semantics while avoiding the significant memory and storage…

Computer Vision and Pattern Recognition · Computer Science 2025-10-13 Yu Sheng , Jiajun Deng , Xinran Zhang , Yu Zhang , Bei Hua , Yanyong Zhang , Jianmin Ji

Language-guided 3D scene understanding is important for advancing applications in robotics, AR/VR, and human-computer interaction, enabling models to comprehend and interact with 3D environments through natural language. While 2D…

Computer Vision and Pattern Recognition · Computer Science 2025-08-05 Anh Thai , Songyou Peng , Kyle Genova , Leonidas Guibas , Thomas Funkhouser

Object navigation is a core capability of embodied intelligence, enabling an agent to locate target objects in unknown environments. Recent advances in vision-language models (VLMs) have facilitated zero-shot object navigation (ZSON).…

Robotics · Computer Science 2026-02-13 Wancai Zheng , Hao Chen , Xianlong Lu , Linlin Ou , Xinyi Yu

Affordance detection and pose estimation are of great importance in many robotic applications. Their combination helps the robot gain an enhanced manipulation capability, in which the generated pose can facilitate the corresponding…

Robotics · Computer Science 2023-09-21 Toan Nguyen , Minh Nhat Vu , Baoru Huang , Tuan Van Vo , Vy Truong , Ngan Le , Thieu Vo , Bac Le , Anh Nguyen

Reconstructing 3D scenes from sparse images remains a challenging task due to the difficulty of recovering accurate geometry and texture without optimization. Recent approaches leverage generalizable models to generate 3D scenes using 3D…

Computer Vision and Pattern Recognition · Computer Science 2026-02-04 Bing He , Jingnan Gao , Yunuo Chen , Ning Cao , Gang Chen , Zhengxue Cheng , Li Song , Wenjun Zhang

Affordance grounding refers to the task of finding the area of an object with which one can interact. It is a fundamental but challenging task, as a successful solution requires the comprehensive understanding of a scene in multiple aspects…

Computer Vision and Pattern Recognition · Computer Science 2024-04-19 Shengyi Qian , Weifeng Chen , Min Bai , Xiong Zhou , Zhuowen Tu , Li Erran Li

Articulated objects, as prevalent entities in human life, their 3D representations play crucial roles across various applications. However, achieving both high-fidelity textured surface reconstruction and dynamic generation for articulated…

Computer Vision and Pattern Recognition · Computer Science 2025-11-24 Di Wu , Liu Liu , Zhou Linli , Anran Huang , Liangtu Song , Qiaojun Yu , Qi Wu , Cewu Lu

We introduce Forecast-aware Gaussian Splatting (Forecast-GS), a predictive 3D representation framework for language-conditioned robotic manipulation. While recent manipulation systems have made progress by grounding language instructions…

Robotics · Computer Science 2026-05-13 Kaixin Jia , Jiacheng Xu

Rendering novel view images in dynamic scenes is a crucial yet challenging task. Current methods mainly utilize NeRF-based methods to represent the static scene and an additional time-variant MLP to model scene deformations, resulting in…

Computer Vision and Pattern Recognition · Computer Science 2024-06-07 Diwen Wan , Ruijie Lu , Gang Zeng

3D Gaussian Splatting (3DGS) has emerged as a novel explicit representation for 3D scenes, offering both high-fidelity reconstruction and efficient rendering. However, 3DGS lacks 3D segmentation ability, which limits its applicability in…

Computer Vision and Pattern Recognition · Computer Science 2025-08-28 Yupeng Zhang , Dezhi Zheng , Ping Lu , Han Zhang , Lei Wang , Liping xiang , Cheng Luo , Kaijun Deng , Xiaowen Fu , Linlin Shen , Jinbao Wang