English
Related papers

Related papers: The Role of Visual Features in Text-Based CAPTCHAs…

200 papers

Supervised image captioning approaches have made great progress, but it is challenging to collect high-quality human-annotated image-text data. Recently, large-scale vision and language models (e.g., CLIP) and large-scale generative…

Computer Vision and Pattern Recognition · Computer Science 2024-03-29 Yiyu Wang , Hao Luo , Jungang Xu , Yingfei Sun , Fan Wang

As a result of continuing advances in computer capabilities, it is becoming increasingly difficult to distinguish between humans and computers in the digital world. We propose using the fundamental human ability to distinguish between…

Cryptography and Security · Computer Science 2019-07-09 Nasser Mohammed Al-Fannah

Understanding the mechanisms underlying human attention is a fundamental challenge for both vision science and artificial intelligence. While numerous computational models of free-viewing have been proposed, less is known about the…

Computer Vision and Pattern Recognition · Computer Science 2023-05-24 Dario Zanca , Andrea Zugarini , Simon Dietz , Thomas R. Altstidl , Mark A. Turban Ndjeuha , Leo Schwinn , Bjoern Eskofier

We propose a new scheme of attack on the Microsoft's ASIRRA CAPTCHA which represents a significant shortcut to the intended attacking path, as it is not based in any advance in the state of the art on the field of image recognition. After…

Cryptography and Security · Computer Science 2009-08-11 Carlos Javier Hernandez-Castro , Arturo Ribagorda , Yago Saez

Content scanning systems employ perceptual hashing algorithms to scan user content for illegal material, such as child pornography or terrorist recruitment flyers. Perceptual hashing algorithms help determine whether two images are visually…

Cryptography and Security · Computer Science 2022-12-09 Ashish Hooda , Andrey Labunets , Tadayoshi Kohno , Earlence Fernandes

Text-based password schemes have inherent security and usability problems, leading to the development of graphical password schemes. However, most of these alternate schemes are vulnerable to spyware attacks. We propose a new scheme, using…

Cryptography and Security · Computer Science 2016-11-18 Liming Wang , Xiuling Chang , Zhongjie Ren , Haichang Gao , Xiyang Liu , Uwe Aickelin

Completely Automated Public Turing tests to tell Computers and Humans Apart (CAPTCHAs) are a foundational component of web security, yet traditional implementations suffer from a trade-off between usability and resilience against AI-powered…

Cryptography and Security · Computer Science 2025-10-06 Ayda Aghaei Nia

Visual modifications to text are often used to obfuscate offensive comments in social media (e.g., "!d10t") or as a writing style ("1337" in "leet speak"), among other scenarios. We consider this as a new type of adversarial attack in NLP,…

Interface icons are prevalent in various digital applications. Due to limited time and budgets, many designers rely on informal evaluation, which often results in poor usability icons. In this paper, we propose a unique human-in-the-loop…

Graphics · Computer Science 2023-05-30 I-Chao Shen , Fu-Yin Cherng , Takeo Igarashi , Wen-Chieh Lin , Bing-Yu Chen

As multimodal large language models (LLMs) advance, traditional CAPTCHAs have become obsolete at distinguishing humans from bots. To address this shift, this paper aims to investigate the possibility of using tasks for which humans have…

Cryptography and Security · Computer Science 2026-04-07 Choon-Hou Rafael Chong

CAPTCHAs are widely used by websites to block bots and spam by presenting challenges that are easy for humans but difficult for automated programs to solve. To improve accessibility, audio CAPTCHAs are designed to complement visual ones.…

Sound · Computer Science 2026-01-14 Ziqi Ding , Yunfeng Wan , Wei Song , Yi Liu , Gelei Deng , Nan Sun , Huadong Mo , Jingling Xue , Shidong Pan , Yuekang Li

The traditional image captioning task uses generic reference captions to provide textual information about images. Different user populations, however, will care about different visual aspects of images. In this paper, we propose a new…

Computation and Language · Computer Science 2020-11-10 Adam Fisch , Kenton Lee , Ming-Wei Chang , Jonathan H. Clark , Regina Barzilay

Image captioning is the process of generating a natural language description of an image. Most current image captioning models, however, do not take into account the emotional aspect of an image, which is very relevant to activities and…

Computer Vision and Pattern Recognition · Computer Science 2019-01-28 Omid Mohamad Nezami , Mark Dras , Peter Anderson , Len Hamey

The rapid evolution of GUI-enabled agents has rendered traditional CAPTCHAs obsolete. While previous benchmarks like OpenCaptchaWorld established a baseline for evaluating multimodal agents, recent advancements in reasoning-heavy models,…

Machine Learning · Computer Science 2026-02-10 Jiacheng Liu , Yaxin Luo , Jiacheng Cui , Xinyi Shang , Xiaohan Zhao , Zhiqiang Shen

In this paper we study a brand new topic of interactive image captioning with human in the loop. Different from automated image captioning where a given test image is the sole input in the inference stage, we have access to both the test…

Human-Computer Interaction · Computer Science 2020-02-25 Zhengxiong Jia , Xirong Li

With the rise of wearables, haptic interfaces are increasingly favored to communicate information in an ambient manner. Despite this expectation, existing guidelines are developed in studies where the participant's focus is entirely on the…

Human-Computer Interaction · Computer Science 2020-06-04 Nava Haghighi , Nathalie Vladis , Yuanbo Liu , Arvind Satyanarayan

CAPTCHAs or reverse Turing tests are real-time assessments used by programs (or computers) to tell humans and machines apart. This is achieved by assigning and assessing hard AI problems that could only be solved easily by human but not by…

Human-Computer Interaction · Computer Science 2014-02-05 A. K. B. Karunathilake , B. M. D. Balasuriya , R. G. Ragel

A precise understanding of why units in an artificial network respond to certain stimuli would constitute a big step towards explainable artificial intelligence. One widely used approach towards this goal is to visualize unit responses via…

Computer Vision and Pattern Recognition · Computer Science 2021-11-15 Roland S. Zimmermann , Judy Borowski , Robert Geirhos , Matthias Bethge , Thomas S. A. Wallis , Wieland Brendel

GUI agents are rapidly shifting from multi-module pipelines to end-to-end, native vision-language models (VLMs) that perceive raw screenshots and directly interact with digital devices. Despite rapid progress on general GUI tasks, CAPTCHA…

Cryptography and Security · Computer Science 2026-03-26 Yuxi Chen , Haoyu Zhai , Chenkai Wang , Rui Yang , Lingming Zhang , Gang Wang , Huan Zhang

Recent lightweight image captioning models using retrieved data mainly focus on text prompts. However, previous works only utilize the retrieved text as text prompts, and the visual information relies only on the CLIP visual embedding.…

Computer Vision and Pattern Recognition · Computer Science 2025-01-27 Taewhan Kim , Soeun Lee , Si-Woo Kim , Dong-Jin Kim