English
Related papers

Related papers: GAIA: A Global, Multi-modal, Multi-scale Vision-La…

200 papers

The rapid advancement of educational applications, artistic creation, and AI-generated content (AIGC) technologies has substantially increased practical requirements for comprehensive Image Aesthetics Assessment (IAA), particularly…

Computer Vision and Pattern Recognition · Computer Science 2025-08-12 Shuo Cao , Nan Ma , Jiayang Li , Xiaohui Li , Lihao Shao , Kaiwen Zhu , Yu Zhou , Yuandong Pu , Jiarui Wu , Jiaquan Wang , Bo Qu , Wenhai Wang , Yu Qiao , Dajuin Yao , Yihao Liu

Image Quality Assessment (IQA) and Image Aesthetic Assessment (IAA) aim to simulate human subjective perception of image visual quality and aesthetic appeal. Despite distinct learning objectives, they have underlying interconnectedness due…

Computer Vision and Pattern Recognition · Computer Science 2025-07-15 Hantao Zhou , Longxiang Tang , Rui Yang , Guanyi Qin , Yan Zhang , Yutao Li , Xiu Li , Runze Hu , Guangtao Zhai

Recently, the diffusion-based generative paradigm has achieved impressive general image generation capabilities with text prompts due to its accurate distribution modeling and stable training process. However, generating diverse remote…

Image and Video Processing · Electrical Eng. & Systems 2024-10-31 Jialin Luo , Yuanzhi Wang , Ziqi Gu , Yide Qiu , Shuaizhen Yao , Fuyun Wang , Chunyan Xu , Wenhua Zhang , Dan Wang , Zhen Cui

Conversational AI tools that can generate and discuss clinically correct radiology reports for a given medical image have the potential to transform radiology. Such a human-in-the-loop radiology assistant could facilitate a collaborative…

Computer Vision and Pattern Recognition · Computer Science 2025-05-08 Chantal Pellegrini , Ege Özsoy , Benjamin Busam , Nassir Navab , Matthias Keicher

The integration of medical images with clinical context is essential for generating accurate and clinically interpretable radiology reports. However, current automated methods often rely on resource-heavy Large Language Models (LLMs) or…

Computer Vision and Pattern Recognition · Computer Science 2025-11-11 Nagur Shareef Shaik , Teja Krishna Cherukuri , Adnan Masood , Dong Hye Ye

Multimodal large language models (MLLMs) have achieved impressive performance across various tasks such as image captioning and visual question answer(VQA); however, they often struggle to accurately interpret depth information inherent in…

Computer Vision and Pattern Recognition · Computer Science 2026-03-09 Hao Yang , Hongbo Zhang , Yanyan Zhao , Bing Qin

Large Vision-Language Models offer a new paradigm for AI-driven image understanding, enabling models to perform tasks without task-specific training. This flexibility holds particular promise across medicine, where expert-annotated data is…

Computer Vision and Pattern Recognition · Computer Science 2025-04-04 Anita Rau , Mark Endo , Josiah Aklilu , Jaewoo Heo , Khaled Saab , Alberto Paderno , Jeffrey Jopling , F. Christopher Holsinger , Serena Yeung-Levy

The recent generative AI models' capability of creating realistic and human-like content is significantly transforming the ways in which people communicate, create and work. The machine-generated content is a double-edged sword. On one…

Computer Vision and Pattern Recognition · Computer Science 2025-02-19 Liting Huang , Zhihao Zhang , Yiran Zhang , Xiyue Zhou , Shoujin Wang

Southeast Asia (SEA) is a region of extraordinary linguistic and cultural diversity, yet it remains significantly underrepresented in vision-language (VL) research. This often results in artificial intelligence (AI) models that fail to…

Computer Vision and Pattern Recognition · Computer Science 2025-03-19 Samuel Cahyawijaya , Holy Lovenia , Joel Ruben Antony Moniz , Tack Hwa Wong , Mohammad Rifqi Farhansyah , Thant Thiri Maung , Frederikus Hudi , David Anugraha , Muhammad Ravi Shulthan Habibi , Muhammad Reza Qorib , Amit Agarwal , Joseph Marvin Imperial , Hitesh Laxmichand Patel , Vicky Feliren , Bahrul Ilmi Nasution , Manuel Antonio Rufino , Genta Indra Winata , Rian Adam Rajagede , Carlos Rafael Catalan , Mohamed Fazli Imam , Priyaranjan Pattnayak , Salsabila Zahirah Pranida , Kevin Pratama , Yeshil Bangera , Adisai Na-Thalang , Patricia Nicole Monderin , Yueqi Song , Christian Simon , Lynnette Hui Xian Ng , Richardy Lobo' Sapan , Taki Hasan Rafi , Bin Wang , Supryadi , Kanyakorn Veerakanjana , Piyalitt Ittichaiwong , Matthew Theodore Roque , Karissa Vincentio , Takdanai Kreangphet , Phakphum Artkaew , Kadek Hendrawan Palgunadi , Yanzhi Yu , Rochana Prih Hastuti , William Nixon , Mithil Bangera , Adrian Xuan Wei Lim , Aye Hninn Khine , Hanif Muhammad Zhafran , Teddy Ferdinan , Audra Aurora Izzani , Ayushman Singh , Evan , Jauza Akbar Krito , Michael Anugraha , Fenal Ashokbhai Ilasariya , Haochen Li , John Amadeo Daniswara , Filbert Aurelian Tjiaranata , Eryawan Presma Yulianrifat , Can Udomcharoenchaikit , Fadil Risdian Ansori , Mahardika Krisna Ihsani , Giang Nguyen , Anab Maulana Barik , Dan John Velasco , Rifo Ahmad Genadi , Saptarshi Saha , Chengwei Wei , Isaiah Flores , Kenneth Ko Han Chen , Anjela Gail Santos , Wan Shen Lim , Kaung Si Phyo , Tim Santos , Meisyarah Dwiastuti , Jiayun Luo , Jan Christian Blaise Cruz , Ming Shan Hee , Ikhlasul Akmal Hanif , M. Alif Al Hakim , Muhammad Rizky Sya'ban , Kun Kerdthaisong , Lester James V. Miranda , Fajri Koto , Tirana Noor Fatyanosa , Alham Fikri Aji , Jostin Jerico Rosal , Jun Kevin , Robert Wijaya , Onno P. Kampman , Ruochen Zhang , Börje F. Karlsson , Peerat Limkonchotiwat

Semantic retrieval of remote sensing (RS) images is a critical task fundamentally challenged by the \textquote{semantic gap}, the discrepancy between a model's low-level visual features and high-level human concepts. While large…

Computer Vision and Pattern Recognition · Computer Science 2025-12-12 J. Xiao , Y. Guo , X. Zi , K. Thiyagarajan , C. Moreira , M. Prasad

Following on recent advances in large language models (LLMs) and subsequent chat models, a new wave of large vision-language models (LVLMs) has emerged. Such models can incorporate images as input in addition to text, and perform tasks such…

Computers and Society · Computer Science 2024-02-09 Kathleen C. Fraser , Svetlana Kiritchenko

Vision-language models (VLMs) have gained widespread adoption in both industry and academia. In this study, we propose a unified framework for systematically evaluating gender, race, and age biases in VLMs with respect to professions. Our…

Computer Vision and Pattern Recognition · Computer Science 2024-06-18 Ashutosh Sathe , Prachi Jain , Sunayana Sitaram

Vision-Language Models (VLMs) have demonstrated effective perception and reasoning capabilities on general-domain tasks, leading to growing interest in their application to Earth observation. However, a systematic benchmark for…

Computer Vision and Pattern Recognition · Computer Science 2026-03-11 Ronghao Fu , Haoran Liu , Weijie Zhang , Zhiwen Lin , Xiao Yang , Peng Zhang , Bo Yang

Multimodal Large Language Models (MLLMs) have become a powerful tool for integrating visual and textual information. Despite their exceptional performance on visual understanding benchmarks, measuring their ability to reason abstractly…

Computer Vision and Pattern Recognition · Computer Science 2026-02-26 Nilay Yilmaz , Maitreya Patel , Yiran Lawrence Luo , Tejas Gokhale , Chitta Baral , Suren Jayasuriya , Yezhou Yang

The emergence of vision language models (VLMs) bridges the gap between vision and language, enabling multimodal understanding beyond traditional visual-only deep learning models. However, transferring VLMs from the natural image domain to…

Computer Vision and Pattern Recognition · Computer Science 2025-11-12 Boyi Li , Ce Zhang , Richard M. Timmerman , Wenxuan Bao

This paper introduces a novel benchmark dataset designed to evaluate the capabilities of Vision Language Models (VLMs) on tasks that combine visual reasoning with subject-specific background knowledge in the German language. In contrast to…

Artificial Intelligence · Computer Science 2025-06-30 René Peinl , Vincent Tischler

The advent of Large Language Models (LLMs) has significantly reshaped the trajectory of the AI revolution. Nevertheless, these LLMs exhibit a notable limitation, as they are primarily adept at processing textual information. To address this…

Computer Vision and Pattern Recognition · Computer Science 2025-10-15 Akash Ghosh , Arkadeep Acharya , Sriparna Saha , Vinija Jain , Aman Chadha

Detecting temporal changes in geographical landscapes is critical for applications like environmental monitoring and urban planning. While remote sensing data is abundant, existing vision-language models (VLMs) often fail to capture…

Computer Vision and Pattern Recognition · Computer Science 2025-05-23 Hosam Elgendy , Ahmed Sharshar , Ahmed Aboeitta , Yasser Ashraf , Mohsen Guizani

Large Vision and Language Models (LVLMs) have shown strong performance across various vision-language tasks in natural image domains. However, their application to remote sensing (RS) remains underexplored due to significant domain…

Computer Vision and Pattern Recognition · Computer Science 2025-06-30 Sungjune Park , Yeongyun Kim , Se Yeon Kim , Yong Man Ro

Segmentation models can recognize a pre-defined set of objects in images. However, models that can reason over complex user queries that implicitly refer to multiple objects of interest are still in their infancy. Recent advances in…

Artificial Intelligence · Computer Science 2025-05-06 Jerome Quenum , Wen-Han Hsieh , Tsung-Han Wu , Ritwik Gupta , Trevor Darrell , David M. Chan
‹ Prev 1 4 5 6 7 8 10 Next ›