English
Related papers

Related papers: SCAM: A Real-World Typographic Robustness Evaluati…

200 papers

Universal adversarial attacks on aligned multimodal large language models are increasingly reported with attack success rates in the 60-80% range, suggesting the visual modality is highly vulnerable to imperceptible perturbations as a…

Cryptography and Security · Computer Science 2026-05-05 Pang Liu , Yingjie Lao

Vertical text input is commonly encountered in various real-world applications, such as mathematical computations and word-based Sudoku puzzles. While current large language models (LLMs) have excelled in natural language tasks, they remain…

Computation and Language · Computer Science 2025-09-29 Zhecheng Li , Yiwei Wang , Bryan Hooi , Yujun Cai , Zhen Xiong , Nanyun Peng , Kai-wei Chang

Recent works have shown the effectiveness of Large Vision Language Models (VLMs or LVLMs) in image manipulation detection. However, text manipulation detection is largely missing in these studies. We bridge this knowledge gap by analyzing…

Computer Vision and Pattern Recognition · Computer Science 2025-09-15 Vidit Vidit , Pavel Korshunov , Amir Mohammadi , Christophe Ecabert , Ketan Kotwal , Sébastien Marcel

Multimodal large language models (MLLMs), which bridge the gap between audio-visual and natural language processing, achieve state-of-the-art performance on several audio-visual tasks. Despite the superior performance of MLLMs, the scarcity…

Cryptography and Security · Computer Science 2025-06-16 Jinming Wen , Xinyi Wu , Shuai Zhao , Yanhao Jia , Yuwen Li

Adversarial attacks expose a fundamental vulnerability in modern deep vision models by exploiting their dependence on dense, pixel-level representations that are highly sensitive to imperceptible perturbations. Traditional defense…

Computer Vision and Pattern Recognition · Computer Science 2025-11-13 Jingjie He , Weijie Liang , Zihan Shan , Matthew Caesar

Face Anti-Spoofing (FAS) typically depends on a single visual modality when defending against presentation attacks such as print attacks, screen replays, and 3D masks, resulting in limited generalization across devices, environments, and…

Computer Vision and Pattern Recognition · Computer Science 2026-03-03 Honglu Zhang , Zhiqin Fang , Ningning Zhao , Saihui Hou , Long Ma , Renwang Pei , Zhaofeng He

Large language model (LLM)-powered multi-agent systems (MAS) enable agents to communicate and share information, achieving strong performance on complex tasks. However, this communication also creates an attack surface where malicious…

Cryptography and Security · Computer Science 2026-05-05 Lingxi Zhang , Guangtao Zheng , Hanjie Chen

Along with the widespread use of face recognition systems, their vulnerability has become highlighted. While existing face anti-spoofing methods can be generalized between attack types, generic solutions are still challenging due to the…

Computer Vision and Pattern Recognition · Computer Science 2022-12-09 Kaicheng Li , Hongyu Yang , Binghui Chen , Pengyu Li , Biao Wang , Di Huang

Today's text-to-image generative models are trained on millions of images sourced from the Internet, each paired with a detailed caption produced by Vision-Language Models (VLMs). This part of the training pipeline is critical for supplying…

Cryptography and Security · Computer Science 2025-06-30 Stanley Wu , Ronik Bhaskar , Anna Yoo Jeong Ha , Shawn Shan , Haitao Zheng , Ben Y. Zhao

Morphing attack detection has become an essential component of face recognition systems for ensuring a reliable verification scenario. In this paper, we present a multimodal learning approach that can provide a textual description of…

Computer Vision and Pattern Recognition · Computer Science 2025-08-15 Sushrut Patwardhan , Raghavendra Ramachandra , Sushma Venkatesh

Visual Question Answering (VQA) is a fundamental task in computer vision and natural language process fields. Although the ``pre-training & finetuning'' learning paradigm significantly improves the VQA performance, the adversarial…

Computer Vision and Pattern Recognition · Computer Science 2024-02-20 Ziyi Yin , Muchao Ye , Tianrong Zhang , Jiaqi Wang , Han Liu , Jinghui Chen , Ting Wang , Fenglong Ma

Vision-language models, which integrate computer vision and natural language processing capabilities, have demonstrated significant advancements in tasks such as image captioning and visual question and answering. However, similar to…

Computer Vision and Pattern Recognition · Computer Science 2025-06-09 Ashwin Ramesh Babu , Sajad Mousavi , Vineet Gundecha , Sahand Ghorbanpour , Avisek Naug , Antonio Guillen , Ricardo Luna Gutierrez , Soumyendu Sarkar

Misinformation can be countered with fact-checking, but the process is costly and slow. Identifying checkworthy claims is the first step, where automation can help scale fact-checkers' efforts. However, detection methods struggle with…

Artificial Intelligence · Computer Science 2025-06-05 Michiel van der Meer , Pavel Korshunov , Sébastien Marcel , Lonneke van der Plas

This paper considers security risks buried in the data processing pipeline in common deep learning applications. Deep learning models usually assume a fixed scale for their training and input data. To allow deep learning applications to…

Cryptography and Security · Computer Science 2017-12-22 Qixue Xiao , Kang Li , Deyue Zhang , Yier Jin

Recent progress in generative AI, primarily through diffusion models, presents significant challenges for real-world deepfake detection. The increased realism in image details, diverse content, and widespread accessibility to the general…

Computer Vision and Pattern Recognition · Computer Science 2024-04-03 Chaitali Bhattacharyya , Hanxiao Wang , Feng Zhang , Sungho Kim , Xiatian Zhu

Recently, the newly emerged multimodal models, which leverage both visual and linguistic modalities to train powerful encoders, have gained increasing attention. However, learning from a large-scale unlabeled dataset also exposes the model…

Cryptography and Security · Computer Science 2023-06-06 Ziqing Yang , Xinlei He , Zheng Li , Michael Backes , Mathias Humbert , Pascal Berrang , Yang Zhang

Modern text-to-image (T2I) models can now render legible, paragraph-length text, enabling a fundamentally new class of misuse. We identify and formalize the inscriptive jailbreak, where an adversary coerces a T2I system into generating…

Computer Vision and Pattern Recognition · Computer Science 2026-04-09 Zonghao Ying , Haowen Dai , Lianyu Hu , Zonglei Jing , Quanchen Zou , Yaodong Yang , Aishan Liu , Xianglong Liu

Recent advances in Multimodal Large Language Models (MLLMs) have shown impressive reasoning capabilities across vision-language tasks, yet still face the challenge of compute-difficulty mismatch. Through empirical analyses, we identify that…

Machine Learning · Computer Science 2026-03-17 Huijie Guo , Jingyao Wang , Lingyu Si , Jiahuan Zhou , Changwen Zheng , Wenwen Qiang

We present THEMIS, a novel multi-task benchmark designed to comprehensively evaluate multimodal large language models (MLLMs) on visual fraud reasoning within real-world academic scenarios. Compared to existing benchmarks, THEMIS introduces…

The rise in popularity of text-to-image generative artificial intelligence (AI) has attracted widespread public interest. We demonstrate that this technology can be attacked to generate content that subtly manipulates its users. We propose…

Computer Vision and Pattern Recognition · Computer Science 2023-09-06 Jordan Vice , Naveed Akhtar , Richard Hartley , Ajmal Mian