English
Related papers

Related papers: Can ChatGPT Perform Image Splicing Detection? A Pr…

200 papers

This study investigates the spatial reasoning capabilities of vision-language models (VLMs) through Chain-of-Thought (CoT) prompting and reinforcement learning. We begin by evaluating the impact of different prompting strategies and find…

Computer Vision and Pattern Recognition · Computer Science 2025-07-21 Binbin Ji , Siddharth Agrawal , Qiance Tang , Yvonne Wu

In today's visually dominated social media landscape, predicting the perceived credibility of visual content and understanding what drives human judgment are crucial for countering misinformation. However, these tasks are challenging due to…

Computer Vision and Pattern Recognition · Computer Science 2025-04-16 Yilang Peng , Sijia Qian , Yingdan Lu , Cuihua Shen

Retrained large language models (LLMs) have become extensively used across various sub-disciplines of natural language processing (NLP). In NLP, text classification problems have garnered considerable focus, but still faced with some…

Computation and Language · Computer Science 2023-12-05 Zhiqiang Wang , Yiran Pang , Yanbin Lin

Cognitive science research treats visual perception, the ability to understand and make sense of a visual input, as one of the early developmental signs of intelligence. Its TVPS-4 framework categorizes and tests human perception into seven…

Computation and Language · Computer Science 2026-01-23 Samrajnee Ghosh , Naman Agarwal , Hemanshu Garg , Chinmay Mittal , Mausam , Parag Singla

Contrastive pretrained large Vision-Language Models (VLMs) like CLIP have revolutionized visual representation learning by providing good performance on downstream datasets. VLMs are 0-shot adapted to a downstream dataset by designing…

Computer Vision and Pattern Recognition · Computer Science 2023-08-09 Mayug Maniparambil , Chris Vorster , Derek Molloy , Noel Murphy , Kevin McGuinness , Noel E. O'Connor

In this paper, we focus on resolving the problem of image outpainting, which aims to extrapolate the surrounding parts given the center contents of an image. Although recent works have achieved promising performance, the lack of versatility…

Computer Vision and Pattern Recognition · Computer Science 2024-12-02 Jinze Yang , Haoran Wang , Zining Zhu , Chenglong Liu , Meng Wymond Wu , Mingming Sun

Recent generative models have demonstrated impressive capabilities in generating realistic and visually pleasing images grounded on textual prompts. Nevertheless, a significant challenge remains in applying these models for the more…

Computer Vision and Pattern Recognition · Computer Science 2025-04-28 Xiaoqian Shen , Mohamed Elhoseiny

The success of large language models (LLMs) has inspired an emerging research field of multimodal learning. However, a grand challenge of exploiting LLMs for multimodal learning is the size of pre-trained LLMs which are always with billions…

Computation and Language · Computer Science 2024-04-08 Zhengqing Yuan , Yunhong He , Kun Wang , Yanfang Ye , Lichao Sun

Patent figure classification facilitates faceted search in patent retrieval systems, enabling efficient prior art search. Existing approaches have explored patent figure classification for only a single aspect and for aspects with a limited…

Information Retrieval · Computer Science 2025-01-23 Sushil Awale , Eric Müller-Budack , Ralph Ewerth

We present GLIPv2, a grounded VL understanding model, that serves both localization tasks (e.g., object detection, instance segmentation) and Vision-Language (VL) understanding tasks (e.g., VQA, image captioning). GLIPv2 elegantly unifies…

Computer Vision and Pattern Recognition · Computer Science 2022-10-13 Haotian Zhang , Pengchuan Zhang , Xiaowei Hu , Yen-Chun Chen , Liunian Harold Li , Xiyang Dai , Lijuan Wang , Lu Yuan , Jenq-Neng Hwang , Jianfeng Gao

This paper introduces MiniGPT4-Video, a multimodal Large Language Model (LLM) designed specifically for video understanding. The model is capable of processing both temporal visual and textual data, making it adept at understanding the…

Computer Vision and Pattern Recognition · Computer Science 2024-04-05 Kirolos Ataallah , Xiaoqian Shen , Eslam Abdelrahman , Essam Sleiman , Deyao Zhu , Jian Ding , Mohamed Elhoseiny

Vision-language models (VLMs) are impactful in part because they can be applied to a variety of visual understanding tasks in a zero-shot fashion, without any fine-tuning. We study $\textit{generative VLMs}$ that are trained for next-word…

Computer Vision and Pattern Recognition · Computer Science 2024-05-16 Zhiqiu Lin , Xinyue Chen , Deepak Pathak , Pengchuan Zhang , Deva Ramanan

Multimodal large language models (MLLMs) that integrate visual and textual reasoning leverage chain-of-thought (CoT) prompting to tackle complex visual tasks, yet continue to exhibit visual hallucinations and an over-reliance on textual…

Computer Vision and Pattern Recognition · Computer Science 2025-10-24 Jing Bi , Guangyu Sun , Ali Vosoughi , Chen Chen , Chenliang Xu

Recent advances in multimodal large language models enable new possibilities for image-based decision support. However, their reliability and operational trade-offs in neuroimaging remain insufficiently understood. We present a…

Computer Vision and Pattern Recognition · Computer Science 2026-03-27 Katarina Trojachanec Dineva , Stefan Andonov , Ilinka Ivanoska , Ivan Kitanovski , Sasho Gramatikov , Tamara Kostova , Monika Simjanoska Misheva , Kostadin Mishev

An emerging family of language models (LMs), capable of processing both text and images within a single visual view, has the promise to unlock complex tasks such as chart understanding and UI navigation. We refer to these models as…

Computation and Language · Computer Science 2024-02-27 Tianyu Gao , Zirui Wang , Adithya Bhaskar , Danqi Chen

Multi-modal Large Language Models (MLLMs) exhibit impressive problem-solving abilities in various domains, but their visual comprehension and abstract reasoning skills remain under-evaluated. To this end, we present PolyMATH, a challenging…

Artificial Intelligence · Computer Science 2026-05-12 Himanshu Gupta , Shreyas Verma , Ujjwala Anantheswaran , Kevin Scaria , Mihir Parmar , Swaroop Mishra , Chitta Baral

In the context of Synthetic Aperture Radar (SAR) image recognition, traditional methods often struggle with the intrinsic limitations of SAR data, such as weak texture, high noise, and ambiguous object boundaries. This work explores a novel…

Signal Processing · Electrical Eng. & Systems 2025-07-15 Chaoran Li , Xingguo Xu , Siyuan Mu

The integration of Large Vision-Language Models (LVLMs) such as OpenAI's GPT-4 Vision into various sectors has marked a significant evolution in the field of artificial intelligence, particularly in the analysis and interpretation of visual…

Computer Vision and Pattern Recognition · Computer Science 2024-12-23 Ahmet Bahaddin Ersoz

$ $The synergy of language and vision models has given rise to Large Language and Vision Assistant models (LLVAs), designed to engage users in rich conversational experiences intertwined with image-based queries. These comprehensive…

Computer Vision and Pattern Recognition · Computer Science 2024-01-02 Ashhadul Islam , Md. Rafiul Biswas , Wajdi Zaghouani , Samir Brahim Belhaouari , Zubair Shah

While Multimodal Large Language Models (MLLMs) have experienced significant advancement in visual understanding and reasoning, their potential to serve as powerful, flexible, interpretable, and text-driven models for Image Quality…

Computer Vision and Pattern Recognition · Computer Science 2024-07-12 Tianhe Wu , Kede Ma , Jie Liang , Yujiu Yang , Lei Zhang