English
Related papers

Related papers: KitchenTwin: Semantically and Geometrically Ground…

200 papers

Multimodal large language models (MLLMs) have exhibited remarkable performance in various visual tasks, yet still struggle with spatial reasoning. Recent efforts mitigate this by injecting geometric features from 3D foundation models, but…

Computer Vision and Pattern Recognition · Computer Science 2026-04-15 Zhaochen Liu , Limeng Qiao , Guanglu Wan , Tingting Jiang

Despite the recent progress on 6D object pose estimation methods for robotic grasping, a substantial performance gap persists between the capabilities of these methods on existing datasets and their efficacy in real-world grasping and…

Robotics · Computer Science 2024-12-18 Abdelrahman Younes , Tamim Asfour

In recent years, the complexity of 5G and beyond wireless networks has escalated, prompting a need for innovative frameworks to facilitate flexible management and efficient deployment. The concept of digital twins (DTs) has emerged as a…

Networking and Internet Architecture · Computer Science 2024-04-24 Zifan Zhang , Mingzhe Chen , Zhaohui Yang , Yuchen Liu

How to extract significant point cloud features and estimate the pose between them remains a challenging question, due to the inherent lack of structure and ambiguous order permutation of point clouds. Despite significant improvements in…

Computer Vision and Pattern Recognition · Computer Science 2021-12-14 Zhu Xu , Zhengyao Bai , Huijie Liu , Qianjie Lu , Shenglan Fan

We present SpatialMem, a memory-centric system for long-horizon, language-grounded retrieval and QA from egocentric video, where metric 3D serves as an interpretable indexing scaffold rather than an explicit mapping objective. Starting from…

Computer Vision and Pattern Recognition · Computer Science 2026-03-09 Xinyi Zheng , Yunze Liu , Chi-Hao Wu , Fan Zhang , Hao Zheng , Wenqi Zhou , Walterio W. Mayol-Cuevas , Junxiao Shen

Quantitative comparison of the quality of photoacoustic image reconstruction algorithms remains a major challenge. No-reference image quality measures are often inadequate, but full-reference measures require access to an ideal reference…

Fringe projection profilometry (FPP) is a widely used technique for measuring object surface form and three-dimensional (3D) geometry, capable of delivering high-precision, high-resolution measurements when paired with suitable cameras and…

Optics · Physics 2026-05-19 D. Weston , X. Kong , G. S. D. Gordon , S. Piano

Video diffusion models generate high-quality and diverse worlds; however, individual frames often lack 3D consistency across the output sequence, which makes the reconstruction of 3D worlds difficult. To this end, we propose a new method…

Computer Vision and Pattern Recognition · Computer Science 2026-03-19 Lukas Höllein , Matthias Nießner

Merging large language models (LLMs) is a practical way to compose capabilities from multiple fine-tuned checkpoints without retraining. Yet standard schemes (linear weight soups, task vectors, and Fisher-weighted averaging) can preserve…

Artificial Intelligence · Computer Science 2025-12-19 Aniruddha Roy , Jyoti Patel , Aman Chadha , Vinija Jain , Amitava Das

This paper presents iMatcher, a fully differentiable framework for feature matching in point cloud registration. The proposed method leverages learned features to predict a geometrically consistent confidence matrix, incorporating both…

Computer Vision and Pattern Recognition · Computer Science 2025-09-12 Karim Slimani , Catherine Achard , Brahim Tamadazte

Vision-language models have been key to the development of open-vocabulary 2D semantic segmentation. Lifting these models from 2D images to 3D scenes, however, remains a challenging problem. Existing approaches typically back-project and…

Computer Vision and Pattern Recognition · Computer Science 2026-04-15 Tomas Berriel Martins , Martin R. Oswald , Javier Civera

Multimodal 3D grounding has garnered considerable interest in Vision-Language Models (VLMs) \cite{yin2025spatial} for advancing spatial reasoning in complex environments. However, these models suffer from a severe "2D semantic bias" that…

Computer Vision and Pattern Recognition · Computer Science 2025-10-21 Yutong Zhong

This paper reports on a novel nonparametric rigid point cloud registration framework that jointly integrates geometric and semantic measurements such as color or semantic labels into the alignment process and does not require explicit data…

Computer Vision and Pattern Recognition · Computer Science 2020-12-08 Ray Zhang , Tzu-Yuan Lin , Chien Erh Lin , Steven A. Parkison , William Clark , Jessy W. Grizzle , Ryan M. Eustice , Maani Ghaffari

Recent volumetric 3D reconstruction methods can produce very accurate results, with plausible geometry even for unobserved surfaces. However, they face an undesirable trade-off when it comes to multi-view fusion. They can fuse all available…

Computer Vision and Pattern Recognition · Computer Science 2021-12-02 Noah Stier , Alexander Rich , Pradeep Sen , Tobias Höllerer

Generalizing metric monocular depth estimation presents a significant challenge due to its ill-posed nature, while the entanglement between camera parameters and depth amplifies issues further, hindering multi-dataset training and zero-shot…

Computer Vision and Pattern Recognition · Computer Science 2025-09-26 Karlo Koledić , Luka Petrović , Ivan Marković , Ivan Petrović

Insufficient data volume and quality are particularly pressing challenges in the adoption of modern subsymbolic AI. To alleviate these challenges, AI simulation uses virtual training environments in which AI agents can be safely and…

Artificial Intelligence · Computer Science 2025-09-01 Xiaoran Liu , Istvan David

Masked image modeling (MIM) has emerged as a promising approach for pre-training Vision Transformers (ViTs). MIMs predict masked tokens token-wise to recover target signals that are tokenized from images or generated by pre-trained models…

Computer Vision and Pattern Recognition · Computer Science 2025-03-24 Taekyung Kim , Byeongho Heo , Dongyoon Han

We present GR3D, a spatial vision language model equipped with three complementary grounding capabilities--explicit 2D grounding, implicit 2D grounding, and monocular 3D grounding--within a single framework. GR3D introduces an implicit…

Computer Vision and Pattern Recognition · Computer Science 2026-05-29 An-Chieh Cheng , Yang Fu , Yatai Ji , Ligeng Zhu , Guanqi Zhan , Zhuoyang Zhang , Zhaojing Yang , Song Han , Yao Lu , Pavlo Molchanov , Vidya Nariyambut Murali , Jan Kautz , Xiaolong Wang , Hongxu Yin , Sifei Liu

LiDAR-based 3D detection has made great progress in recent years. However, the performance of 3D detectors is considerably limited when deployed in unseen environments, owing to the severe domain gap problem. Existing domain adaptive 3D…

Computer Vision and Pattern Recognition · Computer Science 2023-08-17 Ziyu Li , Jingming Guo , Tongtong Cao , Liu Bingbing , Wankou Yang

We propose a new method for fusing a LIDAR point cloud and camera-captured images in the deep convolutional neural network (CNN). The proposed method constructs a new layer called non-homogeneous pooling layer to transform features between…

Computer Vision and Pattern Recognition · Computer Science 2018-02-15 Zining Wang , Wei Zhan , Masayoshi Tomizuka