English
Related papers

Related papers: Granular Privacy Control for Geolocation with Visi…

200 papers

Vision language models (VLMs) can flexibly address various vision tasks through text interactions. Although successful in semantic understanding, state-of-the-art VLMs including GPT-5 still struggle in understanding 3D from 2D inputs. On…

Computer Vision and Pattern Recognition · Computer Science 2025-10-02 Zhipeng Cai , Ching-Feng Yeh , Hu Xu , Zhuang Liu , Gregory Meyer , Xinjie Lei , Changsheng Zhao , Shang-Wen Li , Vikas Chandra , Yangyang Shi

Large language models (LLMs) have shown remarkable capabilities across a broad range of tasks involving question answering and the generation of coherent text and code. Comprehensively understanding the strengths and weaknesses of LLMs is…

Computation and Language · Computer Science 2023-06-02 Jonathan Roberts , Timo Lüddecke , Sowmen Das , Kai Han , Samuel Albanie

Vision language models (VLMs) are designed to extract relevant visuospatial information from images. Some research suggests that VLMs can exhibit humanlike scene understanding, while other investigations reveal difficulties in their ability…

Computer Vision and Pattern Recognition · Computer Science 2025-04-23 Sangeet Khemlani , Tyler Tran , Nathaniel Gyory , Anthony M. Harrison , Wallace E. Lawson , Ravenna Thielstrom , Hunter Thompson , Taaren Singh , J. Gregory Trafton

Recent advancements in multimodal techniques open exciting possibilities for models excelling in diverse tasks involving text, audio, and image processing. Models like GPT-4V, blending computer vision and language modeling, excel in complex…

Computation and Language · Computer Science 2023-10-20 Xiang Zhang , Senyu Li , Zijun Wu , Ning Shi

Multimodal Large Language Models (MLLMs) have demonstrated remarkable capabilities in vision-language tasks. However, these models often infer and reveal sensitive biometric attributes such as race, gender, age, body weight, and eye color;…

Computer Vision and Pattern Recognition · Computer Science 2025-10-08 Younggun Kim , Sirnam Swetha , Fazil Kagdi , Mubarak Shah

Recent advances in visual-language alignment have endowed vision-language models (VLMs) with fine-grained image understanding capabilities. However, this progress also introduces new privacy risks. This paper first proposes a novel privacy…

Computer Vision and Pattern Recognition · Computer Science 2026-03-26 Hongyi Miao , Jun Jia , Xincheng Wang , Qianli Ma , Wei Sun , Wangqiu Zhou , Dandan Zhu , Yewen Cao , Zhi Liu , Guangtao Zhai

Large Vision-Language Models (LVLMs) exhibit impressive potential across various tasks but also face significant privacy risks, limiting their practical applications. Current researches on privacy assessment for LVLMs is limited in scope,…

Cryptography and Security · Computer Science 2026-03-03 Jie Zhang , Xiangkui Cao , Zhouyu Han , Shiguang Shan , Xilin Chen

Large vision-language models (VLMs) such as GPT-4 have achieved unprecedented performance in response generation, especially with visual inputs, enabling more creative and adaptable interaction than large language models such as ChatGPT.…

Computer Vision and Pattern Recognition · Computer Science 2023-10-31 Yunqing Zhao , Tianyu Pang , Chao Du , Xiao Yang , Chongxuan Li , Ngai-Man Cheung , Min Lin

Large Language Model-based Vision-Language Models (LLM-based VLMs) have demonstrated impressive results in various vision-language understanding tasks. However, how well these VLMs can see image detail beyond the semantic level remains…

Computer Vision and Pattern Recognition · Computer Science 2024-08-08 Chenhui Gou , Abdulwahab Felemban , Faizan Farooq Khan , Deyao Zhu , Jianfei Cai , Hamid Rezatofighi , Mohamed Elhoseiny

The recent surge of Multimodal Large Language Models (MLLMs) has fundamentally reshaped the landscape of AI research and industry, shedding light on a promising path toward the next AI milestone. However, significant challenges remain…

Recent large-scale vision-language models (VLMs) have demonstrated remarkable capabilities in understanding and generating textual descriptions for visual content. However, these models lack an understanding of user-specific concepts. In…

Computer Vision and Pattern Recognition · Computer Science 2024-03-22 Yuval Alaluf , Elad Richardson , Sergey Tulyakov , Kfir Aberman , Daniel Cohen-Or

Large-scale Vision Language Models (LVLMs) exhibit advanced capabilities in tasks that require visual information, including object detection. These capabilities have promising applications in various industrial domains, such as autonomous…

Computer Vision and Pattern Recognition · Computer Science 2025-12-01 Haruki Sakajo , Hiroshi Takato , Hiroshi Tsutsui , Komei Soda , Hidetaka Kamigaito , Taro Watanabe

Vision-Language Models (VLMs) have recently shown remarkable progress in multimodal reasoning, yet their applications in autonomous driving remain limited. In particular, the ability to understand road topology, a key requirement for safe…

Computer Vision and Pattern Recognition · Computer Science 2026-03-04 Xin Chen , Jia He , Maozheng Li , Dongliang Xu , Tianyu Wang , Yixiao Chen , Zhixin Lin , Yue Yao

We demonstrate that vision language models (VLMs) are capable of recognizing the content in audio recordings when given corresponding spectrogram images. Specifically, we instruct VLMs to perform audio classification tasks in a few-shot…

Sound · Computer Science 2024-11-20 Satvik Dixit , Laurie M. Heller , Chris Donahue

The next Point-of-Interest (POI) recommendation task aims to predict users' next destinations based on their historical movement data and plays a key role in location-based services and personalized applications. Accurate next POI…

Information Retrieval · Computer Science 2025-05-21 Zhao Liu , Wei Liu , Huajie Zhu , Jianxing Yu , Jian Yin , Wang-Chien Lee , Shun Wang

Large Language Models (LLMs) inherently carry the biases contained in their training corpora, which can lead to the perpetuation of societal harm. As the impact of these foundation models grows, understanding and evaluating their biases…

Computation and Language · Computer Science 2024-10-08 Rohin Manvi , Samar Khanna , Marshall Burke , David Lobell , Stefano Ermon

Vision Large Language Models (VLMs) combine visual understanding with natural language processing, enabling tasks like image captioning, visual question answering, and video analysis. While VLMs show impressive capabilities across domains…

Computer Vision and Pattern Recognition · Computer Science 2025-06-18 Ahmed Sharshar , Latif U. Khan , Waseem Ullah , Mohsen Guizani

Recent breakthroughs in vision-language models (VLMs) emphasize the necessity of benchmarking human preferences in real-world multimodal interactions. To address this gap, we launched WildVision-Arena (WV-Arena), an online platform that…

Computer Vision and Pattern Recognition · Computer Science 2024-06-18 Yujie Lu , Dongfu Jiang , Wenhu Chen , William Yang Wang , Yejin Choi , Bill Yuchen Lin

Vision-Language Models (VLMs), such as GPT-4V and Llama 3.2 vision, have garnered significant research attention for their ability to leverage Large Language Models (LLMs) in multimodal tasks. However, their potential is constrained by…

Computer Vision and Pattern Recognition · Computer Science 2025-02-19 Mukund Agarwalla , Himanshu Kumar , Raj Dandekar , Rajat Dandekar , Sreedath Panat

Large language models (LLMs), such as ChatGPT/GPT-4, have proven to be powerful tools in promoting the user experience as an AI assistant. The continuous works are proposing multi-modal large language models (MLLM), empowering LLMs with the…

Computation and Language · Computer Science 2023-10-23 Ziqiang Zheng , Jipeng Zhang , Tuan-Anh Vu , Shizhe Diao , Yue Him Wong Tim , Sai-Kit Yeung