English
Related papers

Related papers: Tell2Reg: Establishing spatial correspondence betw…

200 papers

Large Vision--Language Models (LVLMs) hold great promise for advancing optical remote sensing (RS) analysis, yet existing reasoning segmentation frameworks couple linguistic reasoning and pixel prediction through end-to-end supervised…

Computer Vision and Pattern Recognition · Computer Science 2026-04-22 Xu Zhang , Junyao Ge , Yang Zheng , Kaitai Guo , Jimin Liang

Computed tomography (CT) report generation is crucial to assist radiologists in interpreting CT volumes, which can be time-consuming and labor-intensive. Existing methods primarily only consider the global features of the entire volume,…

Computer Vision and Pattern Recognition · Computer Science 2025-05-06 Zhixuan Chen , Yequan Bie , Haibo Jin , Hao Chen

Open-vocabulary segmentation models such as SAM3 perform well across broad categories via text prompting, yet degrade when target classes are visually underrepresented in pretraining or depart from canonical depictions-limitations text…

Computer Vision and Pattern Recognition · Computer Science 2026-05-21 Abderrahmene Boudiaf , Irfan Hussain , Sajid Javed

This paper addresses the challenges of registering two rigid semantic scene graphs, an essential capability when an autonomous agent needs to register its map against a remote agent, or against a prior map. The hand-crafted descriptors in…

Robotics · Computer Science 2025-05-21 Chuhao Liu , Zhijian Qiao , Jieqi Shi , Ke Wang , Peize Liu , Shaojie Shen

Regional prompting, or compositional generation, which enables fine-grained spatial control, has gained increasing attention for its practicality in real-world applications. However, previous methods either introduce additional trainable…

Computer Vision and Pattern Recognition · Computer Science 2024-11-19 Zhennan Chen , Yajie Li , Haofan Wang , Zhibo Chen , Zhengkai Jiang , Jun Li , Qian Wang , Jian Yang , Ying Tai

Learning effective region embeddings from heterogeneous urban data underpins key urban computing tasks (e.g., crime prediction, resource allocation). However, prevailing two-stage methods yield task-agnostic representations, decoupling them…

Artificial Intelligence · Computer Science 2026-02-03 Zitao Guo , Changyang Jiang , Tianhong Zhao , Jinzhou Cao , Genan Dai , Bowen Zhang

Foundation models like the segment anything model require high-quality manual prompts for medical image segmentation, which is time-consuming and requires expertise. SAM and its variants often fail to segment structures in ultrasound (US)…

Computer Vision and Pattern Recognition · Computer Science 2025-06-24 Assefa Seyoum Wahd , Banafshe Felfeliyan , Yuyue Zhou , Shrimanti Ghosh , Adam McArthur , Jiechen Zhang , Jacob L. Jaremko , Abhilash Hareendranathan

Multi-organ medical segmentation is a crucial component of medical image processing, essential for doctors to make accurate diagnoses and develop effective treatment plans. Despite significant progress in this field, current multi-organ…

Image and Video Processing · Electrical Eng. & Systems 2025-07-15 Xinlei Yu , Changmiao Wang , Hui Jin , Ahmed Elazab , Gangyong Jia , Xiang Wan , Changqing Zou , Ruiquan Ge

In this paper, we present a method to utilize 2D-2D point matches between images taken during different image conditions to train a convolutional neural network for semantic segmentation. Enforcing label consistency across the matches makes…

Computer Vision and Pattern Recognition · Computer Science 2019-08-19 Måns Larsson , Erik Stenborg , Lars Hammarstrand , Torsten Sattler , Mark Pollefeys , Fredrik Kahl

Many applications, such as autonomous driving, heavily rely on multi-modal data where spatial alignment between the modalities is required. Most multi-modal registration methods struggle computing the spatial correspondence between the…

Computer Vision and Pattern Recognition · Computer Science 2020-03-19 Moab Arar , Yiftach Ginger , Dov Danon , Ilya Leizerson , Amit Bermano , Daniel Cohen-Or

Recognizing spatial relations and reasoning about them is essential in multiple applications including navigation, direction giving and human-computer interaction in general. Spatial relations between objects can either be explicit --…

Computation and Language · Computer Science 2020-07-21 Soham Dan , Hangfeng He , Dan Roth

The main challenge in learning image-conditioned robotic policies is acquiring a visual representation conducive to low-level control. Due to the high dimensionality of the image space, learning a good visual representation requires a…

Robotics · Computer Science 2024-07-03 Albert Yu , Adeline Foote , Raymond Mooney , Roberto Martín-Martín

One of the key shortcomings in current text-to-image (T2I) models is their inability to consistently generate images which faithfully follow the spatial relationships specified in the text prompt. In this paper, we offer a comprehensive…

Grounding DINO and the Segment Anything Model (SAM) have achieved impressive performance in zero-shot object detection and image segmentation, respectively. Together, they have a great potential to revolutionize applications in zero-shot…

Computer Vision and Pattern Recognition · Computer Science 2024-07-02 Fuseini Mumuni , Alhassan Mumuni

Spatial Transcriptomics (ST) reveals the spatial distribution of gene expression in tissues, offering critical insights into biological processes and disease mechanisms. However, the high cost, limited coverage, and technical complexity of…

Computer Vision and Pattern Recognition · Computer Science 2025-04-22 Yi Niu , Jiashuai Liu , Yingkang Zhan , Jiangbo Shi , Di Zhang , Marika Reinius , Ines Machado , Mireia Crispin-Ortuzar , Jialun Wu , Chen Li , Zeyu Gao

Speech-preserving facial expression manipulation (SPFEM) aims to modify facial emotions while meticulously maintaining the mouth animation associated with spoken content. Current works depend on inaccessible paired training samples for the…

Computer Vision and Pattern Recognition · Computer Science 2026-04-23 Tianshui Chen , Jianman Lin , Zhijing Yang , Chunmei Qing , Guangrun Wang , Liang Lin

Text-to-image diffusion models are now capable of generating images that are often indistinguishable from real images. To generate such images, these models must understand the semantics of the objects they are asked to generate. In this…

Computer Vision and Pattern Recognition · Computer Science 2023-12-29 Eric Hedlin , Gopal Sharma , Shweta Mahajan , Hossam Isack , Abhishek Kar , Andrea Tagliasacchi , Kwang Moo Yi

Machine learning for remote sensing imaging relies on up-to-date and accurate labels for model training and testing. Labelling remote sensing imagery is time and cost intensive, requiring expert analysis. Previous labelling tools rely on…

Computer Vision and Pattern Recognition · Computer Science 2026-01-28 Tulsi Patel , Mark W. Jones , Thomas Redfern

We propose a new spatial memory module and a spatial reasoner for the Visual Grounding (VG) task. The goal of this task is to find a certain object in an image based on a given textual query. Our work focuses on integrating the regions of a…

Computer Vision and Pattern Recognition · Computer Science 2021-05-27 Thierry Deruyttere , Guillem Collell , Marie-Francine Moens

Segmenting objects with complex shapes, such as wires, bicycles, or structural grids, remains a significant challenge for current segmentation models, including the Segment Anything Model (SAM) and its high-quality variant SAM-HQ. These…

Computer Vision and Pattern Recognition · Computer Science 2025-06-09 Luka Vetoshkin , Dmitry Yudin