English
Related papers

Related papers: TrianguLang: Geometry-Aware Semantic Consensus for…

200 papers

Understanding the geometric and semantic structure of environments is essential for embodied navigation and reasoning. Existing semantic mapping methods trade off between explicit geometry and multi-scale semantics, and lack a native…

Computer Vision and Pattern Recognition · Computer Science 2026-05-05 Sixian Zhang , Yiyao Wang , Xinhang Song , Keming Zhang , Zijian Xu , Shuqiang Jiang

Modeling open-vocabulary language fields in 3D is essential for intuitive human-AI interaction and querying within physical environments. State-of-the-art approaches, such as LangSplat, leverage 3D Gaussian Splatting to efficiently…

Computer Vision and Pattern Recognition · Computer Science 2025-10-28 Pranav Saxena

Simultaneous Localization and Mapping (SLAM) is a critical task that enables autonomous vehicles to construct maps and localize themselves in unknown environments. Recent breakthroughs combine SLAM with 3D Gaussian Splatting (3DGS) to…

Hardware Architecture · Computer Science 2025-09-03 Houshu He , Naifeng Jing , Li Jiang , Xiaoyao Liang , Zhuoran Song

Vision-language-action (VLA) models have recently shown strong potential in enabling robots to follow language instructions and execute precise actions. However, most VLAs are built upon vision-language models pretrained solely on 2D data,…

Robotics · Computer Science 2025-10-20 Fuhao Li , Wenxuan Song , Han Zhao , Jingbo Wang , Pengxiang Ding , Donglin Wang , Long Zeng , Haoang Li

LiDAR-based 3D detection has made great progress in recent years. However, the performance of 3D detectors is considerably limited when deployed in unseen environments, owing to the severe domain gap problem. Existing domain adaptive 3D…

Computer Vision and Pattern Recognition · Computer Science 2023-08-17 Ziyu Li , Jingming Guo , Tongtong Cao , Liu Bingbing , Wankou Yang

This paper presents a pose-free, feed-forward 3D Gaussian Splatting (3DGS) framework designed to handle unfavorable input views. A common rendering setup for training feed-forward approaches places a 3D object at the world origin and…

Computer Vision and Pattern Recognition · Computer Science 2025-08-07 Yuki Fujimura , Takahiro Kushida , Kazuya Kitano , Takuya Funatomi , Yasuhiro Mukaigawa

Semantic understanding of 3D scenes is essential for robots to operate effectively and safely in complex environments. Existing methods for semantic scene reconstruction and semantic-aware novel view synthesis often rely on dense multi-view…

Computer Vision and Pattern Recognition · Computer Science 2026-03-04 Sheng Ye , Zhen-Hui Dong , Ruoyu Fan , Tian Lv , Yong-Jin Liu

We consider the problem of novel view synthesis from unposed images in a single feed-forward. Our framework capitalizes on fast speed, scalability, and high-quality 3D reconstruction and view synthesis capabilities of 3DGS, where we further…

Computer Vision and Pattern Recognition · Computer Science 2025-07-25 Sunghwan Hong , Jaewoo Jung , Heeseong Shin , Jisang Han , Jiaolong Yang , Chong Luo , Seungryong Kim

We propose an unsupervised vision-based system to estimate the joint configurations of the robot arm from a sequence of RGB or RGB-D images without knowing the model a priori, and then adapt it to the task of category-independent…

Computer Vision and Pattern Recognition · Computer Science 2020-12-02 Qihao Liu , Weichao Qiu , Weiyao Wang , Gregory D. Hager , Alan L. Yuille

Robust visual localization under a wide range of viewing conditions is a fundamental problem in computer vision. Handling the difficult cases of this problem is not only very challenging but also of high practical relevance, e.g., in the…

Computer Vision and Pattern Recognition · Computer Science 2018-04-17 Johannes L. Schönberger , Marc Pollefeys , Andreas Geiger , Torsten Sattler

Estimating metric relative camera pose from a pair of images is of great importance for 3D reconstruction and localisation. However, conventional two-view pose estimation methods are not metric, with camera translation known only up to a…

Computer Vision and Pattern Recognition · Computer Science 2025-09-18 Yumin Li , Dylan Campbell

Injecting semantics into 3D Gaussian Splatting (3DGS) has recently garnered significant attention. While current approaches typically distill 3D semantic features from 2D foundational models (e.g., CLIP and SAM) to facilitate novel view…

Computer Vision and Pattern Recognition · Computer Science 2025-05-20 Wenbo Zhang , Lu Zhang , Ping Hu , Liqian Ma , Yunzhi Zhuge , Huchuan Lu

Visual localization is crucial for Computer Vision and Augmented Reality (AR) applications, where determining the camera or device's position and orientation is essential to accurately interact with the physical environment. Traditional…

Robotics · Computer Science 2025-01-22 Albert Gassol Puigjaner , Irvin Aloise , Patrik Schmuck

Modern 3D object detection datasets are constrained by narrow class taxonomies and costly manual annotations, limiting their ability to scale to open-world settings. In contrast, 2D vision-language models trained on web-scale image-text…

Computer Vision and Pattern Recognition · Computer Science 2025-07-21 Atharv Goel , Mehar Khurana

Existing end-to-end approaches of robotic manipulation often lack generalization to unseen objects or tasks due to limited data and poor interpretability. While recent Multimodal Large Language Models (MLLMs) demonstrate strong commonsense…

Robotics · Computer Science 2026-03-03 Zilong Xie , Jingyu Gong , Xin Tan , Zhizhong Zhang , Yuan Xie

In unstructured environments, functional dexterous grasping calls for the tight integration of semantic understanding, precise 3D functional localization, and physically interpretable execution. Modular hierarchical methods are more…

Computer Vision and Pattern Recognition · Computer Science 2026-04-15 Fan Yang , Wenrui Chen , Guorun Yan , Ruize Liao , Wanjun Jia , Dongsheng Luo , Jiacheng Lin , Kailun Yang , Zhiyong Li , Yaonan Wang

Current techniques in Visual Simultaneous Localization and Mapping (VSLAM) estimate camera displacement by comparing image features of consecutive scenes. These algorithms depend on scene continuity, hence requires frequent camera inputs.…

Robotics · Computer Science 2024-01-25 Mingyang Li , Yue Ma , Qinru Qiu

Grounding natural language questions to functionally relevant regions in 3D objects -- termed language-driven 3D affordance grounding -- is essential for embodied intelligence and human-AI interaction. Existing methods, while progressing…

Computer Vision and Pattern Recognition · Computer Science 2026-03-19 Dongqiang Gou , Xuming He

Real-world robots localize objects from natural-language instructions while scenes around them keep changing. Yet most of the existing 3D visual grounding (3DVG) method still assumes a reconstructed and up-to-date point cloud, an assumption…

Computer Vision and Pattern Recognition · Computer Science 2025-10-17 Miao Hu , Zhiwei Huang , Tai Wang , Jiangmiao Pang , Dahua Lin , Nanning Zheng , Runsen Xu

This paper introduces MipSLAM, a frequency-aware 3D Gaussian Splatting (3DGS) SLAM framework capable of high-fidelity anti-aliased novel view synthesis and robust pose estimation under varying camera configurations. Existing 3DGS-based SLAM…

Computer Vision and Pattern Recognition · Computer Science 2026-03-24 Yingzhao Li , Yan Li , Shixiong Tian , Yanjie Liu , Lijun Zhao , Gim Hee Lee