English
Related papers

Related papers: Unify, Align and Refine: Multi-Level Semantic Alig…

200 papers

Metal artefact reduction (MAR) techniques aim at removing metal-induced noise from clinical images. In Computed Tomography (CT), supervised deep learning approaches have been shown effective but limited in generalisability, as they mostly…

Computer Vision and Pattern Recognition · Computer Science 2020-04-21 Marta B. M. Ranzini , Irme Groothuis , Kerstin Kläser , M. Jorge Cardoso , Johann Henckel , Sébastien Ourselin , Alister Hart , Marc Modat

Liver segmentation on images acquired using computed tomography (CT) and magnetic resonance imaging (MRI) plays an important role in clinical management of liver diseases. Compared to MRI, CT images of liver are more abundant and readily…

Computer Vision and Pattern Recognition · Computer Science 2022-02-25 Jin Hong , Simon Chun-Ho Yu , Weitian Chen

Remote sensing cross-modal text-image retrieval (RSCTIR) has gained attention for its utility in information mining. However, challenges remain in effectively integrating global and local information due to variations in remote sensing…

Computer Vision and Pattern Recognition · Computer Science 2024-11-25 Zengbao Sun , Ming Zhao , Gaorui Liu , André Kaup

Automatic lip-reading (ALR) aims to automatically transcribe spoken content from a speaker's silent lip motion captured in video. Current mainstream lip-reading approaches only use a single visual encoder to model input videos of a single…

Computer Vision and Pattern Recognition · Computer Science 2024-05-01 He Wang , Pengcheng Guo , Xucheng Wan , Huan Zhou , Lei Xie

With the rapid development of Large Language Models (LLMs), aligning these models with human preferences and values is critical to ensuring ethical and safe applications. However, existing alignment techniques such as RLHF or DPO often…

Computation and Language · Computer Science 2025-08-19 Yang Zhang , Yu Yu , Bo Tang , Yu Zhu , Chuxiong Sun , Wenqiang Wei , Jie Hu , Zipeng Xie , Zhiyu Li , Feiyu Xiong , Edward Chung

We introduce a new class of iterative image reconstruction algorithms for radio interferometry, at the interface of convex optimization and deep learning, inspired by plug-and-play methods. The approach consists in learning a prior image…

Image and Video Processing · Electrical Eng. & Systems 2022-09-28 Matthieu Terris , Arwa Dabbech , Chao Tang , Yves Wiaux

Despite the evolution of deep-learning-based visual-textual processing systems, precise multi-modal matching remains a challenging task. In this work, we tackle the task of cross-modal retrieval through image-sentence matching based on…

Computer Vision and Pattern Recognition · Computer Science 2021-03-03 Nicola Messina , Giuseppe Amato , Andrea Esuli , Fabrizio Falchi , Claudio Gennaro , Stéphane Marchand-Maillet

In facial action unit (AU) recognition tasks, regional feature learning and AU relation modeling are two effective aspects which are worth exploring. However, the limited representation capacity of regional features makes it difficult for…

Computer Vision and Pattern Recognition · Computer Science 2021-02-25 Jingwei Yan , Boyuan Jiang , Jingjing Wang , Qiang Li , Chunmao Wang , Shiliang Pu

The increasing prevalence of lumbar spinal canal stenosis has resulted in a surge of MRI (Magnetic Resonance Imaging), leading to labor-intensive interpretation and significant inter-reader variability, even among expert radiologists. This…

Image and Video Processing · Electrical Eng. & Systems 2025-03-04 Arnesh Batra , Arush Gumber , Anushk Kumar

Recently, medical report generation, which aims to automatically generate a long and coherent descriptive paragraph of a given medical image, has received growing research interests. Different from the general image captioning tasks,…

Image and Video Processing · Electrical Eng. & Systems 2022-03-22 Di You , Fenglin Liu , Shen Ge , Xiaoxia Xie , Jing Zhang , Xian Wu

Fine-grained video classification requires understanding complex spatio-temporal and semantic cues that often exceed the capacity of a single modality. In this paper, we propose a multimodal framework that fuses video, image, and text…

Computer Vision and Pattern Recognition · Computer Science 2025-07-08 Namho Kim , Junhwa Kim

Multimodal image registration (MIR) is a fundamental procedure in many image-guided therapies. Recently, unsupervised learning-based methods have demonstrated promising performance over accuracy and efficiency in deformable image…

Computer Vision and Pattern Recognition · Computer Science 2020-11-13 Zhe Xu , Jiangpeng Yan , Jie Luo , Xiu Li , Jayender Jagadeesan

This paper proposes one of the first clinical applications of multimodal large language models (LLMs) as an assistant for radiologists to check errors in their reports. We created an evaluation dataset from real-world radiology datasets…

Computation and Language · Computer Science 2024-03-05 Jinge Wu , Yunsoo Kim , Eva C. Keller , Jamie Chow , Adam P. Levine , Nikolas Pontikos , Zina Ibrahim , Paul Taylor , Michelle C. Williams , Honghan Wu

Unified remote sensing multimodal models exhibit a pronounced spatial reversal curse: Although they can accurately recognize and describe object locations in images, they often fail to faithfully execute the same spatial relations during…

Computer Vision and Pattern Recognition · Computer Science 2026-01-27 Weiyu Zhang , Yuan Hu , Yong Li , Yu Liu

Universal image restoration (UIR) aims to recover clean images from diverse and unknown degradations using a unified model. Existing UIR methods primarily focus on pixel reconstruction and often lack explicit diagnostic reasoning over…

Computer Vision and Pattern Recognition · Computer Science 2026-04-13 Wending Yan , Rongkai Zhang , Kaihua Tang , Yu Cheng , Qiankun Liu

Multimodal emotion recognition systems rely heavily on the full availability of modalities, suffering significant performance declines when modal data is incomplete. To tackle this issue, we present the Cross-Modal Alignment,…

Multimedia · Computer Science 2024-07-15 Haoqin Sun , Shiwan Zhao , Shaokai Li , Xiangyu Kong , Xuechen Wang , Aobo Kong , Jiaming Zhou , Yong Chen , Wenjia Zeng , Yong Qin

Cross-modal alignment Learning integrates information from different modalities like text, image, audio and video to create unified models. This approach develops shared representations and learns correlations between modalities, enabling…

Computer Vision and Pattern Recognition · Computer Science 2024-09-19 Bilal Faye , Hanane Azzag , Mustapha Lebbah

Low-rank adaptation (LoRA) has emerged as a powerful tool for parameter-efficient fine-tuning of large language models (LLMs). This paper studies LoRA under a federated learning setting, enabling collaborative fine-tuning across clients…

Machine Learning · Statistics 2026-05-21 Shuaida He , Liwen Chen , Long Feng

Recent advancements in image generation models have enabled personalized image creation with both user-defined subjects (content) and styles. Prior works achieved personalization by merging corresponding low-rank adapters (LoRAs) through…

Computer Vision and Pattern Recognition · Computer Science 2025-08-12 Donald Shenaj , Ondrej Bohdal , Mete Ozay , Pietro Zanuttigh , Umberto Michieli

Automated radiology report generation has gained increasing attention with the rise of deep learning and large language models. However, fully generative approaches often suffer from hallucinations and lack clinical grounding, limiting…

Quantitative Methods · Quantitative Biology 2026-05-01 Himadri S Samanta
‹ Prev 1 4 5 6 7 8 10 Next ›