English
Related papers

Related papers: FUSAR-GPT : A Spatiotemporal Feature-Embedded and …

200 papers

We present LatentAM, an online 3D Gaussian Splatting (3DGS) mapping framework that builds scalable latent feature maps from streaming RGB-D observations for open-vocabulary robotic perception. Instead of distilling high-dimensional…

Robotics · Computer Science 2026-02-16 Junwoon Lee , Yulun Tian

Robust and accurate perception of dynamic objects and map elements is crucial for autonomous vehicles performing safe navigation in complex traffic scenarios. While vision-only methods have become the de facto standard due to their…

Speckle noise poses a significant challenge in maintaining the quality of synthetic aperture radar (SAR) images, so SAR despeckling techniques have drawn increasing attention. Despite the tremendous advancements of deep learning in…

Computer Vision and Pattern Recognition · Computer Science 2024-01-09 Xuran Hu , Ziqiang Xu , Zhihan Chen , Zhengpeng Feng , Mingzhe Zhu , LJubisa Stankovic

Novel-view synthesis (NVS) approaches play a critical role in vast scene reconstruction. However, these methods rely heavily on dense image inputs and prolonged training times, making them unsuitable where computational resources are…

Computer Vision and Pattern Recognition · Computer Science 2024-11-21 Hao Li , Yuanyuan Gao , Haosong Peng , Chenming Wu , Weicai Ye , Yufeng Zhan , Chen Zhao , Dingwen Zhang , Jingdong Wang , Junwei Han

Visual Text-to-Speech (VTTS) aims to take the environmental image as the prompt to synthesize the reverberant speech for the spoken content. The challenge of this task lies in understanding the spatial environment from the image. Many…

Computer Vision and Pattern Recognition · Computer Science 2025-01-16 Rui Liu , Shuwei He , Yifan Hu , Haizhou Li

Cross-View Geo-Localization (CVGL) estimates the location of a ground image by matching it to a geo-tagged aerial image in a database. Recent works achieve outstanding progress on CVGL benchmarks. However, existing methods still suffer from…

Computer Vision and Pattern Recognition · Computer Science 2024-08-15 Xiaohan Zhang , Xingyu Li , Waqas Sultani , Chen Chen , Safwan Wshah

Diffusion-based image super-resolution (SR) methods have demonstrated remarkable performance. Recent advancements have introduced deterministic sampling processes that reduce inference from 15 iterative steps to a single step, thereby…

Computer Vision and Pattern Recognition · Computer Science 2025-05-13 Zihang Liu , Zhenyu Zhang , Hao Tang

Large language models (LLMs) are increasingly used in emergency first response (EFR) applications to support situational awareness (SA) and decision-making, yet most operate on text or 2D imagery and offer little support for core EFR SA…

Human-Computer Interaction · Computer Science 2026-02-18 Rodrigo Gutierrez Maquilon , Marita Hueber , Georg Regal , Manfred Tscheligi

Image super-resolution (SR) has significantly advanced through the adoption of Transformer architectures. However, conventional techniques aimed at enlarging the self-attention window to capture broader contexts come with inherent…

Computer Vision and Pattern Recognition · Computer Science 2025-03-20 Chengxing Xie , Xiaoming Zhang , Linze Li , Yuqian Fu , Biao Gong , Tianrui Li , Kai Zhang

Applying Gaussian Splatting to perception tasks for 3D scene understanding is becoming increasingly popular. Most existing works primarily focus on rendering 2D feature maps from novel viewpoints, which leads to an imprecise 3D language…

Computer Vision and Pattern Recognition · Computer Science 2026-03-10 Hao Li , Minghan Qin , Zhengyu Zou , Diqi He , Xinhao Ji , Bohan Li , Bingquan Dai , Dingewn Zhang , Junwei Han

Lip Reading, or Visual Automatic Speech Recognition (V-ASR), is a complex task requiring the interpretation of spoken language exclusively from visual cues, primarily lip movements and facial expressions. This task is especially challenging…

Computer Vision and Pattern Recognition · Computer Science 2026-01-06 Marshall Thomas , Edward Fish , Richard Bowden

Empowered by large-scale training, vision-language models (VLMs) achieve strong image and video understanding, yet their ability to perform spatial reasoning in both static scenes and dynamic videos remains limited. Recent advances try to…

Computer Vision and Pattern Recognition · Computer Science 2026-03-30 Shihua Zhang , Qiuhong Shen , Shizun Wang , Tianbo Pan , Xinchao Wang

Large Vision--Language Models (LVLMs) hold great promise for advancing optical remote sensing (RS) analysis, yet existing reasoning segmentation frameworks couple linguistic reasoning and pixel prediction through end-to-end supervised…

Computer Vision and Pattern Recognition · Computer Science 2026-04-22 Xu Zhang , Junyao Ge , Yang Zheng , Kaitai Guo , Jimin Liang

We introduce GSU, a text-only grid dataset to evaluate the spatial reasoning capabilities of LLMs over 3 core tasks: navigation, object localization, and structure composition. By forgoing visual inputs, isolating spatial reasoning from…

Computation and Language · Computer Science 2026-03-19 Risham Sidhu , Julia Hockenmaier

Neural surface reconstruction (NSR) has recently shown strong potential for urban 3D reconstruction from multi-view aerial imagery. However, existing NSR methods often suffer from geometric ambiguity and instability, particularly under…

Computer Vision and Pattern Recognition · Computer Science 2026-01-30 Da Li , Chen Yao , Tong Mao , Jiacheng Bao , Houjun Sun

Recently, employing single-modality large language models based on mechanical vibration signals as Tuning Predictors has introduced new perspectives in intelligent fault diagnosis. However, the potential of these methods to leverage…

Emerging Technologies · Computer Science 2025-02-24 Jiao Chen , Ruyi Huang , Zuohong Lv , Jianhua Tang , Weihua Li

Decision-makers in GIS need to combine a series of spatial algorithms and operations to solve geospatial tasks. For example, in the task of facility siting, the Buffer tool is usually first used to locate areas close or away from some…

Computation and Language · Computer Science 2023-07-18 Yifan Zhang , Cheng Wei , Shangyou Wu , Zhengting He , Wenhao Yu

Contact-rich robotic manipulation requires representations that encode local geometry. Vision provides global context but lacks direct measurements of properties such as texture and hardness, whereas touch supplies these cues. Modern…

Computer Vision and Pattern Recognition · Computer Science 2025-12-02 Gurmeher Khurana , Lan Wei , Dandan Zhang

Scene understanding based on 3D Gaussian Splatting (3DGS) has recently achieved notable advances. Although 3DGS related methods have efficient rendering capabilities, they fail to address the inherent contradiction between the anisotropic…

Computer Vision and Pattern Recognition · Computer Science 2025-05-27 Q. G. Duan , Benyun Zhao , Mingqiao Han Yijun Huang , Ben M. Chen

The advent of immersive Virtual Reality applications has transformed various domains, yet their integration with advanced artificial intelligence technologies like Visual Language Models remains underexplored. This study introduces a…

Robotics · Computer Science 2024-08-06 Mikhail Konenkov , Artem Lykov , Daria Trinitatova , Dzmitry Tsetserukou