English
Related papers

Related papers: VAP-Diffusion: Enriching Descriptions with MLLMs f…

200 papers

Inpainting focuses on filling missing or corrupted regions of an image to blend seamlessly with its surrounding content and style. While conditional diffusion models have proven effective for text-guided inpainting, we introduce the novel…

Computer Vision and Pattern Recognition · Computer Science 2025-04-11 Nicola Fanelli , Gennaro Vessio , Giovanna Castellano

Medical image segmentation has immense clinical applicability but remains a challenge despite advancements in deep learning. The Segment Anything Model (SAM) exhibits potential in this field, yet the requirement for expertise intervention…

Computer Vision and Pattern Recognition · Computer Science 2023-11-14 Yinsong Xu , Jiaqi Tang , Aidong Men , Qingchao Chen

In autoregressive (AR) image generation, models based on the 'next-token prediction' paradigm of LLMs have shown comparable performance to diffusion models by reducing inductive biases. However, directly applying LLMs to complex image…

Computer Vision and Pattern Recognition · Computer Science 2025-03-13 Miaomiao Cai , Guanjie Wang , Wei Li , Zhijun Tu , Hanting Chen , Shaohui Lin , Jie Hu

Multimodal Large Language Models (MLLMs) show promise for medical applications, yet progress in dermatology lags due to limited training data, narrow task coverage, and lack of clinically-grounded supervision that mirrors expert diagnostic…

Computation and Language · Computer Science 2026-01-06 Jinghan Ru , Siyuan Yan , Yuguo Yin , Yuexian Zou , Zongyuan Ge

Medical Large Vision-Language Models (Med-LVLMs) have been widely adopted for medical report generation. Despite Med-LVLMs producing state-of-the-art performance, they exhibit a bias toward predicting all findings as normal, leading to…

Multiagent Systems · Computer Science 2025-05-27 Pengyu Wang , Shuchang Ye , Usman Naseem , Jinman Kim

We introduce the new task of generating Illustrated Instructions, i.e., visual instructions customized to a user's needs. We identify desiderata unique to this task, and formalize it through a suite of automatic and human evaluation…

Computer Vision and Pattern Recognition · Computer Science 2024-04-16 Sachit Menon , Ishan Misra , Rohit Girdhar

Zero-shot medical detection can further improve detection performance without relying on annotated medical images even upon the fine-tuned model, showing great clinical value. Recent studies leverage grounded vision-language models (GLIP)…

Computer Vision and Pattern Recognition · Computer Science 2025-02-25 Yuguang Yang , Tongfei Chen , Haoyu Huang , Linlin Yang , Chunyu Xie , Dawei Leng , Xianbin Cao , Baochang Zhang

Compared with Large Language Models (LLMs), Large Vision-Language Models (LVLMs) can also accept images as input, thus showcasing more interesting emergent capabilities and demonstrating impressive performance on various vision-language…

Computer Vision and Pattern Recognition · Computer Science 2024-09-26 Runpeng Yu , Weihao Yu , Xinchao Wang

Large vision-language models (VLMs) have advanced multimodal tasks such as video question answering (QA). However, VLMs face the challenge of selecting frames effectively and efficiently, as standard uniform sampling is expensive and…

Computer Vision and Pattern Recognition · Computer Science 2026-05-05 Martin Q. Ma , Willis Guo , Aditya Agrawal , Ankit Gupta , Paul Pu Liang , Ruslan Salakhutdinov , Louis-Philippe Morency

Recent advances in multimodal large language models (MLLMs) and diffusion models (DMs) have opened new possibilities for AI-generated content. Yet, personalized cover image generation remains underexplored, despite its critical role in…

Computation and Language · Computer Science 2026-05-28 Zhipeng Bian , Jieming Zhu , Qijiong Liu , Wang Lin , Guohao Cai , Zhaocheng Du , Jiacheng Sun , Zhou Zhao , Zhenhua Dong

Large Language Models (LLMs) exhibit strong natural language processing capabilities but also inherit and amplify societal biases, including gender bias, raising fairness concerns. Existing debiasing methods face significant limitations:…

Computation and Language · Computer Science 2025-02-18 Hongye Qiu , Yue Xu , Meikang Qiu , Wenjie Wang

3D medical image analysis is pivotal in numerous clinical applications. However, the scarcity of labeled data and limited generalization capabilities hinder the advancement of AI-empowered models. Radiology reports are easily accessible and…

Computer Vision and Pattern Recognition · Computer Science 2024-12-10 Xuefeng Ni , Linshan Wu , Jiaxin Zhuang , Qiong Wang , Mingxiang Wu , Varut Vardhanabhuti , Lihai Zhang , Hanyu Gao , Hao Chen

Generative Large Language Models (LLMs) hold significant promise in healthcare, demonstrating capabilities such as passing medical licensing exams and providing clinical knowledge. However, their current use as information retrieval tools…

Existing image perception methods based on VLMs generally follow a paradigm wherein models extract and analyze image content based on user-provided textual task prompts. However, such methods face limitations when applied to UAV imagery,…

Computer Vision and Pattern Recognition · Computer Science 2025-12-09 Mingning Guo , Mengwei Wu , Shaoxian Li , Haifeng Li , Chao Tao

In medical image analysis, the expertise scarcity and the high cost of data annotation limits the development of large artificial intelligence models. This paper investigates the potential of transfer learning with pre-trained…

Computer Vision and Pattern Recognition · Computer Science 2024-05-27 Jiajin Zhang , Ge Wang , Mannudeep K. Kalra , Pingkun Yan

Vision-language models have emerged as a powerful tool for previously challenging multi-modal classification problem in the medical domain. This development has led to the exploration of automated image description generation for…

Computer Vision and Pattern Recognition · Computer Science 2024-06-03 Mansi Kakkar , Dattesh Shanbhag , Chandan Aladahalli , Gurunath Reddy M

In medical image classification, supervised learning is challenging due to the scarcity of labeled medical images. To address this, we leverage the visual-textual alignment within Vision-Language Models (VLMs) to enable unsupervised…

Computer Vision and Pattern Recognition · Computer Science 2025-04-01 Umaima Rahman , Raza Imam , Mohammad Yaqub , Boulbaba Ben Amor , Dwarikanath Mahapatra

Generating high-fidelity 3D avatars from text or image prompts is highly sought after in virtual reality and human-computer interaction. However, existing text-driven methods often rely on iterative Score Distillation Sampling (SDS) or CLIP…

Computer Vision and Pattern Recognition · Computer Science 2026-03-05 Hong Li , Yutang Feng , Minqi Meng , Yichen Yang , Xuhui Liu , Baochang Zhang

Vision-language foundation models (VLMs) show promise for diverse imaging tasks but often underperform on medical benchmarks. Prior efforts to improve performance include model finetuning, which requires large domain-specific datasets and…

Computer Vision and Pattern Recognition · Computer Science 2025-11-18 Arnav Singhvi , Vasiliki Bikia , Asad Aali , Akshay Chaudhari , Roxana Daneshjou

Histopathology image classification is crucial for the accurate identification and diagnosis of various diseases but requires large and diverse datasets. Obtaining such datasets, however, is often costly and time-consuming due to the need…

Computer Vision and Pattern Recognition · Computer Science 2024-09-25 Leire Benito-Del-Valle , Aitor Alvarez-Gila , Itziar Eguskiza , Cristina L. Saratxaga