English
Related papers

Related papers: AmaraSpatial-10K: A Spatially and Semantically Ali…

200 papers

Recent advances in 3D Gaussian splatting have significantly improved real-time novel view synthesis, yet insufficient geometric constraints during scene optimization often result in blurred reconstructions of fine-grained details,…

Computer Vision and Pattern Recognition · Computer Science 2025-08-15 Zheng Zhou , Jia-Chen Zhang , Yu-Jie Xiong , Chun-Ming Xia

Embodied intelligence fundamentally requires a capability to determine where to act in 3D space. We formalize this requirement as embodied localization -- the problem of predicting executable 3D points conditioned on visual observations and…

Robotics · Computer Science 2026-03-31 Qiming Zhu , Zhirui Fang , Tianming Zhang , Chuanxiu Liu , Xiaoke Jiang , Lei Zhang

Accurate and efficient global ocean state estimation remains a grand challenge for Earth system science, hindered by the dual bottlenecks of computational scalability and degraded data fidelity in traditional data assimilation (DA) and deep…

Machine Learning · Computer Science 2025-11-11 Yanfei Xiang , Yuan Gao , Hao Wu , Quan Zhang , Ruiqi Shu , Xiao Zhou , Xi Wu , Xiaomeng Huang

In this paper, we introduce a novel benchmark designed to propel the advancement of general-purpose, large-scale 3D vision models for remote sensing imagery. While several datasets have been proposed within the realm of remote sensing, many…

Computer Vision and Pattern Recognition · Computer Science 2025-09-24 Jiayu Wang , Ruizhi Wang , Jie Song , Haofei Zhang , Mingli Song , Zunlei Feng , Li Sun

Decomposing 3D assets into material parts is a common task for artists, yet remains a highly manual process. In this work, we introduce Select Any Material (SAMa), a material selection approach for in-the-wild objects in arbitrary 3D…

Computer Vision and Pattern Recognition · Computer Science 2026-02-23 Michael Fischer , Iliyan Georgiev , Thibault Groueix , Vladimir G. Kim , Tobias Ritschel , Valentin Deschaintre

Multimodal large language models (MLLMs) have advanced static visual--spatial reasoning, yet they often fail to preserve long-horizon spatial coherence in embodied settings where beliefs must be continuously revised from egocentric…

Computer Vision and Pattern Recognition · Computer Science 2026-05-29 Chih-Ting Liao , Xi Xiao , Chunlei Meng , Zhangquan Chen , Yitong Qiao , Weilin Zhou , Tianyang Wang , Xu Zheng , Xin Cao

In this paper, we propose a RGB-D SLAM system that reconstructs a language-aligned dense feature field while sustaining low-latency tracking and mapping. First, we introduce a Top-K Rendering pipeline, a high-throughput and…

Robotics · Computer Science 2026-02-10 Seongbo Ha , Sibaek Lee , Kyungsu Kang , Joonyeol Choi , Seungjun Tak , Hyeonwoo Yu

We propose 3DSmoothNet, a full workflow to match 3D point clouds with a siamese deep learning architecture and fully convolutional layers using a voxelized smoothed density value (SDV) representation. The latter is computed per interest…

Computer Vision and Pattern Recognition · Computer Science 2019-12-03 Zan Gojcic , Caifa Zhou , Jan D. Wegner , Andreas Wieser

Many underwater applications, such as offshore asset inspections, rely on visual inspection and detailed 3D reconstruction. Recent advancements in underwater visual SLAM systems for aquatic environments have garnered significant attention…

Robotics · Computer Science 2025-06-06 Yifan Peng , Yuze Hong , Ziyang Hong , Apple Pui-Yi Chui , Junfeng Wu

Better understanding and modelling of building interiors and the emergence of more impressive AR/VR technology has brought up the need for automatic parsing of floorplan images. However, there is a clear lack of representative datasets to…

Computer Vision and Pattern Recognition · Computer Science 2019-04-04 Ahti Kalervo , Juha Ylioinas , Markus Häikiö , Antti Karhu , Juho Kannala

Increasing the volume of training data can enable the auxiliary diagnostic algorithms for Autism Spectrum Disorder (ASD) to learn more accurate and stable models. However, due to the significant heterogeneity and domain shift in rs-fMRI…

Neurons and Cognition · Quantitative Biology 2025-07-11 Yiqian Luo , Qiurong Chen , Fali Li , Peng Xu , Yangsong Zhang

The Segment Anything Model (SAM) has demonstrated impressive generalization in prompt-based segmentation. Yet, the potential of semantic text prompts remains underexplored compared to traditional spatial prompts like points and boxes. This…

Computer Vision and Pattern Recognition · Computer Science 2025-08-04 Shayan Jalilian , Abdul Bais

In industrial applications requiring real-time feedback, such as quality control and robotic manipulation, the demand for high-speed and accurate pose estimation remains critical. Despite advances improving speed and accuracy in pose…

Computer Vision and Pattern Recognition · Computer Science 2025-02-21 Zixuan Fang , Thomas Pöllabauer , Tristan Wirth , Sarah Berkei , Volker Knauthe , Arjan Kuijper

Segment anything models (SAMs) are gaining attention for their zero-shot generalization capability in segmenting objects of unseen classes and in unseen domains when properly prompted. Interactivity is a key strength of SAMs, allowing users…

Image and Video Processing · Electrical Eng. & Systems 2024-03-18 Yiqing Shen , Jingxing Li , Xinyuan Shao , Blanca Inigo Romillo , Ankush Jindal , David Dreizin , Mathias Unberath

Significant progress has been made in spatial intelligence, spanning both spatial reconstruction and world exploration. However, the scalability and real-world fidelity of current models remain severely constrained by the scarcity of…

Computer Vision and Pattern Recognition · Computer Science 2025-12-19 Jiahao Wang , Yufeng Yuan , Rujie Zheng , Youtian Lin , Jian Gao , Lin-Zhuo Chen , Yajie Bao , Yi Zhang , Chang Zeng , Yanxi Zhou , Xiao-Xiao Long , Hao Zhu , Zhaoxiang Zhang , Xun Cao , Yao Yao

Recent advancements in Multimodal Large Language Models (MLLMs) have significantly enhanced performance on 2D visual tasks. However, improving their spatial intelligence remains a challenge. Existing 3D MLLMs always rely on additional 3D or…

Computer Vision and Pattern Recognition · Computer Science 2026-05-20 Diankun Wu , Fangfu Liu , Yi-Hsin Hung , Yueqi Duan

Multi-sensor Simultaneous Localization and Mapping (SLAM) is essential for Unmanned Aerial Vehicles (UAVs) performing agricultural tasks such as spraying, surveying, and inspection. However, real-world, multi-modal agricultural UAV datasets…

Robotics · Computer Science 2026-01-13 Zhihao Zhan , Yuhang Ming , Shaobin Li , Jie Yuan

One of the key shortcomings in current text-to-image (T2I) models is their inability to consistently generate images which faithfully follow the spatial relationships specified in the text prompt. In this paper, we offer a comprehensive…

Accurate 3D point cloud registration underpins reliable image-guided colonoscopy, directly affecting lesion localization, margin assessment, and navigation safety. However, biological tissue exhibits repetitive textures and locally…

Computer Vision and Pattern Recognition · Computer Science 2025-11-04 Linzhe Jiang , Jiayuan Huang , Sophia Bano , Matthew J. Clarkson , Zhehua Mao , Mobarak I. Hoque

While 3D Gaussian Splatting (3DGS) enabled photorealistic mapping, its integration into SLAM has largely followed traditional camera-centric pipelines. As a result, they inherit well-known weaknesses such as high computational load, failure…

Robotics · Computer Science 2026-03-10 Jaeseok Park , Chanoh Park , Minsu Kim , Minkyoung Kim , Soohwan Kim