English
Related papers

Related papers: Large Language Models for Multimodal Deformable Im…

200 papers

Multimodal recommender systems (MRS) integrate heterogeneous user and item data, such as text, images, and structured information, to enhance recommendation performance. The emergence of large language models (LLMs) introduces new…

Information Retrieval · Computer Science 2025-05-16 Alejo Lopez-Avila , Jinhua Du

Large language models (LLMs) have demonstrated remarkable language abilities. GPT-4, based on advanced LLMs, exhibits extraordinary multimodal capabilities beyond previous visual language models. We attribute this to the use of more…

Computation and Language · Computer Science 2023-05-23 Feilong Chen , Minglun Han , Haozhi Zhao , Qingyang Zhang , Jing Shi , Shuang Xu , Bo Xu

Multimodal large language models (MLLMs) have recently become a focal point of research due to their formidable multimodal understanding capabilities. For example, in the audio and speech domains, an LLM can be equipped with (automatic)…

Computer Vision and Pattern Recognition · Computer Science 2025-03-10 Umberto Cappellazzo , Minsu Kim , Honglie Chen , Pingchuan Ma , Stavros Petridis , Daniele Falavigna , Alessio Brutti , Maja Pantic

Ensuring fairness across demographic groups in medical diagnosis is essential for equitable healthcare, particularly under distribution shifts caused by variations in imaging equipment and clinical practice. Vision-language models (VLMs)…

Computer Vision and Pattern Recognition · Computer Science 2025-08-27 Yuexuan Xia , Benteng Ma , Jiang He , Zhiyong Wang , Qi Dou , Yong Xia

In-context learning (ICL) facilitates Large Language Models (LLMs) exhibiting emergent ability on downstream tasks without updating billions of parameters. However, in the area of multi-modal Large Language Models (MLLMs), two problems…

Multimedia · Computer Science 2024-07-02 Jun Gao , Qian Qiao , Ziqiang Cao , Zili Wang , Wenjie Li

Text-to-Image Person Retrieval (TIPR) is a cross-modal matching task designed to identify the person images that best correspond to a given textual description. The key difficulty in TIPR is to realize robust correspondence between the…

Computer Vision and Pattern Recognition · Computer Science 2025-12-30 Hao Yin , Xin Man , Feiyu Chen , Jie Shao , Heng Tao Shen

Metal-organic frameworks (MOFs) are porous crystalline materials with broad applications such as carbon capture and drug delivery, yet accurately predicting their 3D structures remains a significant challenge. While Large Language Models…

Machine Learning · Computer Science 2026-01-15 Mianzhi Pan , JianFei Li , Peishuo Liu , Botian Wang , Yawen Ouyang , Yiming Rong , Hao Zhou , Jianbing Zhang

Heterogeneous face recognition (HFR) refers to matching face images acquired from different domains with wide applications in security scenarios. This paper presents a deep neural network approach namely Multi-Margin based Decorrelation…

Computer Vision and Pattern Recognition · Computer Science 2020-05-26 Bing Cao , Nannan Wang , Xinbo Gao , Jie Li , Zhifeng Li

We have seen remarkable success in representation learning and language models (LMs) using deep neural networks. Many studies aim to build the underlying connections among different modalities via the alignment and mappings at the token or…

Sound · Computer Science 2025-03-04 Daniel Chin , Gus Xia

The advancement of Multimodal Large Language Models (MLLMs) has greatly accelerated the development of applications in understanding integrated texts and images. Recent works leverage image-caption datasets to train MLLMs, achieving…

Computation and Language · Computer Science 2024-11-22 Mingxu Tao , Quzhe Huang , Kun Xu , Liwei Chen , Yansong Feng , Dongyan Zhao

Current medical image segmentation approaches have limitations in deeply exploring multi-scale information and effectively combining local detail textures with global contextual semantic information. This results in over-segmentation,…

Computer Vision and Pattern Recognition · Computer Science 2025-01-08 Zhenkun Lu , Chaoyin She , Wei Wang , Qinghua Huang

The primary challenges in visible-infrared person re-identification arise from the differences between visible (vis) and infrared (ir) images, including inter-modal and intra-modal variations. These challenges are further complicated by…

Computer Vision and Pattern Recognition · Computer Science 2024-09-10 Jiarui Li , Zhen Qiu , Yilin Yang , Yuqi Li , Zeyu Dong , Chuanguang Yang

The emergence of Multimodal Large Language Models (MLLMs) has revolutionized image understanding by bridging textual and visual modalities. However, these models often struggle with capturing fine-grained semantic information, such as the…

Computer Vision and Pattern Recognition · Computer Science 2025-07-16 Jie Yang , Wang Zeng , Sheng Jin , Lumin Xu , Wentao Liu , Chen Qian , Zhen Li , Ruimao Zhang

The field of object detection and understanding is rapidly evolving, driven by advances in both traditional CNN-based models and emerging multi-modal large language models (LLMs). While CNNs like ResNet and YOLO remain highly effective for…

Computer Vision and Pattern Recognition · Computer Science 2025-10-13 Nirmal Elamon , Rouzbeh Davoudi

Designing generative models for 3D structural brain MRI that synthesize morphologically-plausible and attribute-specific (e.g., age, sex, disease state) samples is an active area of research. Existing approaches based on frameworks like…

Image and Video Processing · Electrical Eng. & Systems 2025-08-04 Alan Q. Wang , Fangrui Huang , Bailey Trang , Wei Peng , Mohammad Abbasi , Kilian Pohl , Mert Sabuncu , Ehsan Adeli

Unsupervised deep learning is a promising method in brain MRI registration to reduce the reliance on anatomical labels, while still achieving anatomically accurate transformations. For the Learn2Reg2024 LUMIR challenge, we propose…

Image and Video Processing · Electrical Eng. & Systems 2024-12-31 Lukas Förner , Kartikay Tehlan , Thomas Wendler

Large language models (LLMs) are primarily designed to understand unstructured text. When directly applied to structured formats such as tabular data, they may struggle to discern inherent relationships and overlook critical patterns. While…

Machine Learning · Computer Science 2024-10-11 Natraj Raman , Sumitra Ganesh , Manuela Veloso

The visual commonsense reasoning (VCR) task is to choose an answer and provide a justifying rationale based on the given image and textural question. Representative works first recognize objects in images and then associate them with key…

Computer Vision and Pattern Recognition · Computer Science 2023-12-27 Jian Zhu , Hanli Wang , Miaojing Shi

Existing deep learning-based methods can capture shared features from optical and synthetic aperture radar (SAR) images for spatial alignment. However, optical-SAR registration remains challenging under large geometric deformations, because…

Computer Vision and Pattern Recognition · Computer Science 2026-05-13 Zhuoyu Cai , Dou Quan , Ning Huyan , Pei He , Shuang Wang , Licheng Jiao

Large language models (LLMs) excel in diverse applications but face dual challenges: generating harmful content under jailbreak attacks and over-refusal of benign queries due to rigid safety mechanisms. These issues are further complicated…

Artificial Intelligence · Computer Science 2025-11-04 Yifan Xia , Guorui Chen , Wenqian Yu , Zhijiang Li , Philip Torr , Jindong Gu