English
Related papers

Related papers: SCAM: A Real-World Typographic Robustness Evaluati…

200 papers

The proliferation of synthetic images generated by advanced AI models poses significant challenges in identifying and understanding manipulated visual content. Current fake image detection methods predominantly rely on binary classification…

Computer Vision and Pattern Recognition · Computer Science 2025-03-20 Ritabrata Chakraborty , Rajatsubhra Chakraborty , Ali Khaleghi Rahimian , Thomas MacDougall

The rapid development of audio-driven talking head generators and advanced Text-To-Speech (TTS) models has led to more sophisticated temporal deepfakes. These advances highlight the need for robust methods capable of detecting and…

Audio and Speech Processing · Electrical Eng. & Systems 2025-08-12 Ivan Kukanov , Jun Wah Ng

Face Anti-Spoofing (FAS) is essential for ensuring the security and reliability of facial recognition systems. Most existing FAS methods are formulated as binary classification tasks, providing confidence scores without interpretation. They…

Computer Vision and Pattern Recognition · Computer Science 2025-01-27 Guosheng Zhang , Keyao Wang , Haixiao Yue , Ajian Liu , Gang Zhang , Kun Yao , Errui Ding , Jingdong Wang

Recent advances in text-to-image diffusion models have enabled the generation of diverse and high-quality images. While impressive, the images often fall short of depicting subtle details and are susceptible to errors due to ambiguity in…

Computer Vision and Pattern Recognition · Computer Science 2025-01-13 Idan Schwartz , Vésteinn Snæbjarnarson , Hila Chefer , Ryan Cotterell , Serge Belongie , Lior Wolf , Sagie Benaim

Research on Multi-modal Large Language Models (MLLMs) towards the multi-image cross-modal instruction has received increasing attention and made significant progress, particularly in scenarios involving closely resembling images (e.g.,…

Computer Vision and Pattern Recognition · Computer Science 2024-08-26 Tao Wu , Mengze Li , Jingyuan Chen , Wei Ji , Wang Lin , Jinyang Gao , Kun Kuang , Zhou Zhao , Fei Wu

Multimodal Large Language Models (MLLMs) demonstrate remarkable capabilities that increasingly influence various aspects of our daily lives, constantly defining the new boundary of Artificial General Intelligence (AGI). Image modalities,…

Cryptography and Security · Computer Science 2024-08-13 Yihe Fan , Yuxin Cao , Ziyu Zhao , Ziyao Liu , Shaofeng Li

While neural machine translation (NMT) models achieve success in our daily lives, they show vulnerability to adversarial attacks. Despite being harmful, these attacks also offer benefits for interpreting and enhancing NMT models, thus…

Computation and Language · Computer Science 2024-09-10 Yanni Xue , Haojie Hao , Jiakai Wang , Qiang Sheng , Renshuai Tao , Yu Liang , Pu Feng , Xianglong Liu

Deep visual models are susceptible to adversarial perturbations to inputs. Although these signals are carefully crafted, they still appear noise-like patterns to humans. This observation has led to the argument that deep visual…

Computer Vision and Pattern Recognition · Computer Science 2021-06-22 Naveed Akhtar , Muhammad A. A. K. Jalwana , Mohammed Bennamoun , Ajmal Mian

With the advent of Large Language Models (LLMs) possessing increasingly impressive capabilities, a number of Large Vision-Language Models (LVLMs) have been proposed to augment LLMs with visual inputs. Such models condition generated text on…

Computer Vision and Pattern Recognition · Computer Science 2025-05-01 Phillip Howard , Kathleen C. Fraser , Anahita Bhiwandiwalla , Svetlana Kiritchenko

Deep learning models achieve remarkable accuracy in computer vision tasks, yet remain vulnerable to adversarial examples--carefully crafted perturbations to input images that can deceive these models into making confident but incorrect…

Computer Vision and Pattern Recognition · Computer Science 2025-04-18 Khoi Nguyen Tiet Nguyen , Wenyu Zhang , Kangkang Lu , Yuhuan Wu , Xingjian Zheng , Hui Li Tan , Liangli Zhen

Textual interaction networks (TINs) are an omnipresent data structure used to model the interplay between users and items on e-commerce websites, social networks, etc., where each interaction is associated with a text description.…

Computation and Language · Computer Science 2025-04-08 Hongtao Wang , Renchi Yang , Hewen Wang , Haoran Zheng , Jianliang Xu

Multimodal large language models (MLLMs) have demonstrated strong capabilities in vision-related tasks, capitalizing on their visual semantic comprehension and reasoning capabilities. However, their ability to detect subtle visual spoofing…

Computer Vision and Pattern Recognition · Computer Science 2025-04-22 Yichen Shi , Yuhao Gao , Yingxin Lai , Hongyang Wang , Jun Feng , Lei He , Jun Wan , Changsheng Chen , Zitong Yu , Xiaochun Cao

Phishing websites now rely heavily on visual imitation-copied logos, similar layouts, and matching colours-to avoid detection by text- and URL-based systems. This paper presents a deep learning approach that uses webpage screenshots for…

Computer Vision and Pattern Recognition · Computer Science 2026-04-16 K. Acharya , S. Ale , R. Kadel

With the rapid advancement of generative AI, multimodal deepfakes, which manipulate both audio and visual modalities, have drawn increasing public concern. Currently, deepfake detection has emerged as a crucial strategy in countering these…

Sound · Computer Science 2024-05-16 Yang Hou , Haitao Fu , Chuankai Chen , Zida Li , Haoyu Zhang , Jianjun Zhao

Recent research on large language models (LLMs) has demonstrated their ability to understand and employ deceptive behavior, even without explicit prompting. However, such behavior has only been observed in rare, specialized cases and has…

Computation and Language · Computer Science 2025-06-24 Laurène Vaugrante , Francesca Carlon , Maluna Menke , Thilo Hagendorff

AI-generated synthetic media are increasingly used in real-world scenarios, often with the purpose of spreading misinformation and propaganda through social media platforms, where compression and other processing can degrade fake detection…

Multimedia · Computer Science 2025-04-30 Stefano Dell'Anna , Andrea Montibeller , Giulia Boato

Synthetic image generation has opened up new opportunities but has also created threats in regard to privacy, authenticity, and security. Detecting fake images is of paramount importance to prevent illegal activities, and previous research…

Computer Vision and Pattern Recognition · Computer Science 2023-02-27 Md Awsafur Rahman , Bishmoy Paul , Najibul Haque Sarker , Zaber Ibn Abdul Hakim , Shaikh Anowarul Fattah

Large-scale vision models have become integral in many applications due to their unprecedented performance and versatility across downstream tasks. However, the robustness of these foundation models has primarily been explored for a single…

Computer Vision and Pattern Recognition · Computer Science 2024-07-19 Antoni Kowalczuk , Jan Dubiński , Atiyeh Ashari Ghomi , Yi Sui , George Stein , Jiapeng Wu , Jesse C. Cresswell , Franziska Boenisch , Adam Dziedzic

Unstructured text from legal, medical, and administrative sources offers a rich but underutilized resource for research in public health and the social sciences. However, large-scale analysis is hampered by two key challenges: the presence…

Computation and Language · Computer Science 2025-07-16 Anders Ledberg , Anna Thalén

Vision-Language Models (VLMs) have gained considerable prominence in recent years due to their remarkable capability to effectively integrate and process both textual and visual information. This integration has significantly enhanced…

Computer Vision and Pattern Recognition · Computer Science 2025-02-12 Aobotao Dai , Xinyu Ma , Lei Chen , Songze Li , Lin Wang