中文
相关论文

相关论文: Indescribable Multi-modal Spatial Evaluator

200 篇论文

A challenge in high-dimensional inverse problems is developing iterative solvers to find the accurate solution of regularized optimization problems with low computational cost. An important example is computed tomography (CT) where both…

数值分析 · 数学 2024-12-16 Alessandro Perelli , Carola-Bibiane Schonlieb , Matthias J. Ehrhardt

Robust 3D registration is a fundamental problem in computer vision and robotics, where the goal is to estimate the geometric transformation between two sets of measurements in the presence of noise, mismatches, and extreme outlier…

计算机视觉与模式识别 · 计算机科学 2026-04-09 Xianyun Qian , Fei Wen , Peilin Liu

Photoacoustic tomography (PAT) offers optical contrast, whereas magnetic resonance imaging (MRI) excels in imaging soft tissue and organ anatomy. The fusion of PAT with MRI holds promising application prospects due to their complementary…

图像与视频处理 · 电气工程与系统科学 2025-03-20 Yutian Zhong , Jinchuan He , Zhichao Liang , Shuangyang Zhang , Qianjin Feng , Lijun Lu , Li Qi

End-to-end In-Image Machine Translation (IIMT) aims to convert text embedded within an image into a target language while preserving the original visual context, layout, and rendering style. However, existing IIMT benchmarks are largely…

计算机视觉与模式识别 · 计算机科学 2026-04-02 Jiahao Lyu , Pei Fu , Zhenhang Li , Weichao Zeng , Shaojie Zhang , Jiahui Yang , Can Ma , Yu Zhou , Zhenbo Luo , Jian Luan

This work investigates the problem of instance-level image retrieval re-ranking with the constraint of memory efficiency, ultimately aiming to limit memory usage to 1KB per image. Departing from the prevalent focus on performance…

计算机视觉与模式识别 · 计算机科学 2024-08-07 Pavel Suma , Giorgos Kordopatis-Zilos , Ahmet Iscen , Giorgos Tolias

Anomaly detection and localization in medical imaging remain critical challenges in healthcare. This paper introduces Spatial-MSMA (Multiscale Score Matching Analysis), a novel unsupervised method for anomaly localization in volumetric…

计算机视觉与模式识别 · 计算机科学 2024-07-19 Ahsan Mahmood , Junier Oliva , Martin Styner

Ensemble smoother (ES) has been widely used in inverse modeling of hydrologic systems. However, for problems where the distribution of model parameters is multimodal, using ES directly would be problematic. One popular solution is to use a…

最优化与控制 · 数学 2018-02-27 Jiangjiang Zhang , Guang Lin , Weixuan Li , Laosheng Wu , Lingzao Zeng

The success of many computer vision tasks lies in the ability to exploit the interdependency between different image modalities such as intensity and depth. Fusing corresponding information can be achieved on several levels, and one…

计算机视觉与模式识别 · 计算机科学 2014-06-26 Martin Kiechle , Tim Habigt , Simon Hawe , Martin Kleinsteuber

Very recently, Window-based Transformers, which computed self-attention within non-overlapping local windows, demonstrated promising results on image classification, semantic segmentation, and object detection. However, less study has been…

计算机视觉与模式识别 · 计算机科学 2021-06-08 Zilong Huang , Youcheng Ben , Guozhong Luo , Pei Cheng , Gang Yu , Bin Fu

Image registration is a fundamental medical image analysis task. Ideally, registration should focus on aligning semantically corresponding voxels, i.e., the same anatomical locations. However, existing methods often optimize similarity…

计算机视觉与模式识别 · 计算机科学 2024-02-27 Lin Tian , Zi Li , Fengze Liu , Xiaoyu Bai , Jia Ge , Le Lu , Marc Niethammer , Xianghua Ye , Ke Yan , Daikai Jin

Multimodal image-text models have shown remarkable performance in the past few years. However, evaluating robustness against distribution shifts is crucial before adopting them in real-world applications. In this work, we investigate the…

计算机视觉与模式识别 · 计算机科学 2024-01-22 Jielin Qiu , Yi Zhu , Xingjian Shi , Florian Wenzel , Zhiqiang Tang , Ding Zhao , Bo Li , Mu Li

Multi-modal learning adeptly integrates visual and textual data, but its application to histopathology image and text analysis remains challenging, particularly with large, high-resolution images like gigapixel Whole Slide Images (WSIs).…

计算机视觉与模式识别 · 计算机科学 2024-05-29 Quan Liu , Ruining Deng , Can Cui , Tianyuan Yao , Vishwesh Nath , Yucheng Tang , Yuankai Huo

The field of advanced text-to-image generation is witnessing the emergence of unified frameworks that integrate powerful text encoders, such as CLIP and T5, with Diffusion Transformer backbones. Although there have been efforts to control…

计算机视觉与模式识别 · 计算机科学 2025-02-28 Liang Chen , Shuai Bai , Wenhao Chai , Weichu Xie , Haozhe Zhao , Leon Vinci , Junyang Lin , Baobao Chang

Deformable image registration remains a central challenge in medical image analysis, particularly under multi-modal scenarios where intensity distributions vary significantly across scans. While deep learning methods provide efficient…

图像与视频处理 · 电气工程与系统科学 2026-03-30 Yi Zhang , Yidong Zhao , Qian Tao

Multi-focus image fusion aims to combine multiple partially focused images into a single all-in-focus image. Although deep learning has shown promise in this task, its effectiveness is often limited by the scarcity of suitable training…

计算机视觉与模式识别 · 计算机科学 2026-03-10 Huangxing Lin , Rongrong Ma , Cheng Wang

Most of existing manifold learning methods rely on Mean Squared Error (MSE) or $\ell_2$ norm. However, for the problem of image quality assessment, these are not promising measure. In this paper, we introduce the concept of an image…

机器学习 · 统计学 2019-08-27 Benyamin Ghojogh , Fakhri Karray , Mark Crowley

We propose a coercive approach to simultaneously register and segment multi-modal images which share similar spatial structure. Registration is done at the region level to facilitate data fusion while avoiding the need for interpolation.…

计算机视觉与模式识别 · 计算机科学 2015-11-19 Yu-Hui Chen , Dennis Wei , Gregory Newstadt , Jeffrey Simmons , Alfred Hero

In most scenarios, conditional image generation can be thought of as an inversion of the image understanding process. Since generic image understanding involves solving multiple tasks, it is natural to aim at generating images via…

计算机视觉与模式识别 · 计算机科学 2022-07-15 Ritika Chakraborty , Nikola Popovic , Danda Pani Paudel , Thomas Probst , Luc Van Gool

Establishing dense correspondences between multiple images is a fundamental task in many applications. However, finding a reliable correspondence in multi-modal or multi-spectral images still remains unsolved due to their challenging…

计算机视觉与模式识别 · 计算机科学 2016-04-28 Seungryong Kim , Dongbo Min , Bumsub Ham , Minh N. Do , Kwanghoon Sohn

In this work, we present a novel, machine-learning approach for constructing Multiclass Interpretable Scoring Systems (MISS) - a fully data-driven methodology for generating single, sparse, and user-friendly scoring systems for multiclass…

机器学习 · 计算机科学 2024-01-11 Michal K. Grzeszczyk , Tomasz Trzciński , Arkadiusz Sitek