中文
相关论文

相关论文: Tell2Reg: Establishing spatial correspondence betw…

200 篇论文

Training embodied agents in simulation has become mainstream for the embodied AI community. However, these agents often struggle when deployed in the physical world due to their inability to generalize to real-world environments. In this…

机器人学 · 计算机科学 2022-12-12 Matt Deitke , Rose Hendrix , Luca Weihs , Ali Farhadi , Kiana Ehsani , Aniruddha Kembhavi

Non-rigid inter-modality registration can facilitate accurate information fusion from different modalities, but it is challenging due to the very different image appearances across modalities. In this paper, we propose to train a non-rigid…

计算机视觉与模式识别 · 计算机科学 2018-05-01 Xiaohuan Cao , Jianhua Yang , Li Wang , Zhong Xue , Qian Wang , Dinggang Shen

Prompt engineering has shown remarkable success with large language models, yet its systematic exploration in computer vision remains limited. In semantic segmentation, both textual and visual prompts offer distinct advantages: textual…

计算机视觉与模式识别 · 计算机科学 2025-08-07 Gabriele Rosi , Fabio Cermelli

Large Vision-Language Models (VLMs) are increasingly being regarded as foundation models that can be instructed to solve diverse tasks by prompting, without task-specific training. We examine the seemingly obvious question: how to…

计算机视觉与模式识别 · 计算机科学 2026-02-09 Niccolo Avogaro , Thomas Frick , Mattia Rigotti , Andrea Bartezzaghi , Filip Janicki , Cristiano Malossi , Konrad Schindler , Roy Assaf

Arctic Permafrost is facing significant changes due to global climate change. As these regions are largely inaccessible, remote sensing plays a crucial rule in better understanding the underlying processes not just on a local scale, but…

计算机视觉与模式识别 · 计算机科学 2024-01-18 Konrad Heidler , Ingmar Nitze , Guido Grosse , Xiao Xiang Zhu

We present a method for finding cross-modal space-time correspondences. Given two images from different visual modalities, such as an RGB image and a depth map, our model identifies which pairs of pixels correspond to the same physical…

计算机视觉与模式识别 · 计算机科学 2025-06-04 Ayush Shrivastava , Andrew Owens

Mobile robots necessitate advanced natural language understanding capabilities to accurately identify locations and perform tasks such as package delivery. However, traditional visual place recognition (VPR) methods rely solely on…

计算机视觉与模式识别 · 计算机科学 2025-03-10 Tianyi Shang , Zhenyu Li , Pengjie Xu , Jinwei Qiao , Gang Chen , Zihan Ruan , Weijun Hu

Multi-modality medical images can provide relevant or complementary information for a target (organ, tumor or tissue). Registering multi-modality images to a common space can fuse these comprehensive information, and bring convenience for…

图像与视频处理 · 电气工程与系统科学 2021-08-31 Wangbin Ding , Lei Li , Xiahai Zhuang , Liqin Huang

Researchers in functional neuroimaging mostly use activation coordinates to formulate their hypotheses. Instead, we propose to use the full statistical images to define regions of interest (ROIs). This paper presents two machine learning…

机器学习 · 统计学 2012-09-10 Yannick Schwartz , Gaël Varoquaux , Bertrand Thirion

An estimated half of the world's languages do not have a written form, making it impossible for these languages to benefit from any existing text-based technologies. In this paper, a speech-to-image generation (S2IG) framework is proposed…

机器学习 · 计算机科学 2020-09-16 Xinsheng Wang , Tingting Qiao , Jihua Zhu , Alan Hanjalic , Odette Scharenborg

Text-to-image (T2I) diffusion models generate high-quality images but often fail to capture the spatial relations specified in text prompts. This limitation can be traced to two factors: lack of fine-grained spatial supervision in training…

计算机视觉与模式识别 · 计算机科学 2025-12-30 Sarah Rastegar , Violeta Chatalbasheva , Sieger Falkena , Anuj Singh , Yanbo Wang , Tejas Gokhale , Hamid Palangi , Hadi Jamali-Rad

Estimating correspondence between two images and extracting the foreground object are two challenges in computer vision. With dual-lens smart phones, such as iPhone 7Plus and Huawei P9, coming into the market, two images of slightly…

计算机视觉与模式识别 · 计算机科学 2017-04-10 Xiaoyong Shen , Hongyun Gao , Xin Tao , Chao Zhou , Jiaya Jia

Learning latent representations that capture both semantic and spatial information is central to efficient spatio-semantic reasoning. However, many existing approaches rely on implicit latent structures combined with dense feature maps or…

计算机视觉与模式识别 · 计算机科学 2026-05-13 SeongMin Jin , Doo Seok Jeong

Earth observation offers new insight into anthropogenic changes to nature, and how these changes are effecting (and are effected by) the built environment and the real economy. With the global availability of medium-resolution (10-30m)…

计算机视觉与模式识别 · 计算机科学 2021-02-15 Lucas Kruitwagen

Referring remote sensing image segmentation is crucial for achieving fine-grained visual understanding through free-format textual input, enabling enhanced scene and object extraction in remote sensing applications. Current research…

计算机视觉与模式识别 · 计算机科学 2025-01-14 Keyan Chen , Jiafan Zhang , Chenyang Liu , Zhengxia Zou , Zhenwei Shi

Advances in machine learning, especially the introduction of transformer architectures and vision transformers, have led to the development of highly capable computer vision foundation models. The segment anything model (known colloquially…

计算机视觉与模式识别 · 计算机科学 2025-05-29 Kenneth Ball , Erin Taylor , Nirav Patel , Andrew Bartels , Gary Koplik , James Polly , Jay Hineman

Spatial Transcriptomics is a novel technology that aligns histology images with spatially resolved gene expression profiles. Although groundbreaking, it struggles with gene capture yielding high corruption in acquired data. Given potential…

计算机视觉与模式识别 · 计算机科学 2024-09-30 Gabriel Mejia , Daniela Ruiz , Paula Cárdenas , Leonardo Manrique , Daniela Vega , Pablo Arbeláez

Visual Geo-localization (VG) refers to the process to identify the location described in query images, which is widely applied in robotics field and computer vision tasks, such as autonomous driving, metaverse, augmented reality, and SLAM.…

计算机视觉与模式识别 · 计算机科学 2024-06-05 Chen Mao , Jingqi Hu

Human environments contain numerous objects configured in a variety of arrangements. Our goal is to enable robots to repose previously unseen objects according to learned semantic relationships in novel environments. We break this problem…

机器人学 · 计算机科学 2021-08-30 Chris Paxton , Chris Xie , Tucker Hermans , Dieter Fox

Existing audio-driven visual dubbing methods have achieved great success. Despite this, we observe that the semantic ambiguity between spatial and temporal domains significantly degrades the synthesis stability for the dynamic faces. We…

计算机视觉与模式识别 · 计算机科学 2025-04-01 Zijun Ding , Mingdie Xiong , Congcong Zhu , Jingrun Chen