English
Related papers

Related papers: VeriDark: A Large-Scale Benchmark for Authorship V…

200 papers

Authorship identification tasks, which rely heavily on linguistic styles, have always been an important part of Natural Language Understanding (NLU) research. While other tasks based on linguistic style understanding benefit from deep…

Computation and Language · Computer Science 2020-10-01 Weicheng Ma , Ruibo Liu , Lili Wang , Soroush Vosoughi

One of the most challenging forms of misinformation involves pairing images with misleading text to create false narratives. Existing AI-driven detection systems often require domain-specific finetuning, limiting generalizability, and offer…

Artificial Intelligence · Computer Science 2025-10-07 Kumud Lakara , Georgia Channing , Christian Rupprecht , Juil Sock , Philip Torr , John Collomosse , Christian Schroeder de Witt

In this paper, we describe DeFactoNLP, the system we designed for the FEVER 2018 Shared Task. The aim of this task was to conceive a system that can not only automatically assess the veracity of a claim but also retrieve evidence supporting…

Artificial Intelligence · Computer Science 2018-09-10 Aniketh Janardhan Reddy , Gil Rocha , Diego Esteves

Deep learning based face-swap videos, widely known as deepfakes, have drawn wide attention due to their threat to information credibility. Recent works mainly focus on the problem of deepfake detection that aims to reliably tell deepfakes…

Computer Vision and Pattern Recognition · Computer Science 2026-03-20 Bo Peng , Zichuan Wang , Sheng Yu , Xiaochuan Jin , Wei Wang , Jing Dong

The hidden nature and the limited accessibility of the Dark Web, combined with the lack of public datasets in this domain, make it difficult to study its inherent characteristics such as linguistic properties. Previous works on text…

Computation and Language · Computer Science 2022-05-05 Youngjin Jin , Eugene Jang , Yongjae Lee , Seungwon Shin , Jin-Woo Chung

In the field of fraud detection, the availability of comprehensive and privacy-compliant datasets is crucial for advancing machine learning research and developing effective anti-fraud systems. Traditional datasets often focus on…

Machine Learning · Computer Science 2024-04-24 Phoebe Jing , Yijing Gao , Xianlong Zeng

Application systems using natural language interfaces to databases (NLIDBs) have democratized data analysis. This positive development has also brought forth an urgent challenge to help users who might use these systems without a background…

Computation and Language · Computer Science 2025-07-25 Shubham Mohole , Sainyam Galhotra

With the spread of online social networks, it is more and more difficult to monitor all the user-generated content. Automating the moderation process of the inappropriate exchange content on Internet has thus become a priority task. Methods…

Computation and Language · Computer Science 2021-01-19 Noé Cecillon , Vincent Labatut , Richard Dufour , Georges Linares

Image repurposing is a commonly used method for spreading misinformation on social media and online forums, which involves publishing untampered images with modified metadata to create rumors and further propaganda. While manual…

Computer Vision and Pattern Recognition · Computer Science 2019-04-11 Ayush Jaiswal , Yue Wu , Wael AbdAlmageed , Iacopo Masi , Premkumar Natarajan

With the growth of fake news and disinformation, the NLP community has been working to assist humans in fact-checking. However, most academic research has focused on model accuracy without paying attention to resource efficiency, which is…

Computers and Society · Computer Science 2021-09-03 Mykola Trokhymovych , Diego Saez-Trumper

Underground forums where users discuss, buy, and sell illicit services and goods facilitate a better understanding of the economy and organization of cybercriminals. Prior work has shown that in particular private interactions provide a…

Cryptography and Security · Computer Science 2018-05-14 Rebekah Overdorf , Carmela Troncoso , Rachel Greenstadt , Damon McCoy

The proliferation of AIGC-driven face manipulation and deepfakes poses severe threats to media provenance, integrity, and copyright protection. Prior versatile watermarking systems typically rely on embedding explicit localization payloads,…

Computer Vision and Pattern Recognition · Computer Science 2026-03-26 Peipeng Yu , Jinfeng Xie , Chengfu Ou , Xiaoyu Zhou , Jianwei Fei , Yunshu Dai , Zhihua Xia , Chip Hong Chang

Vulnerability detection has always been the most important task in the field of software security. With the development of technology, in the face of massive source code, automated analysis and detection of vulnerabilities has become a…

Cryptography and Security · Computer Science 2021-04-26 Jiajie Wu

Recent measurements of the Windows code-signing certificate ecosystem have highlighted various forms of abuse that allow malware authors to produce malicious code carrying valid digital signatures. However, the underground trade that allows…

Cryptography and Security · Computer Science 2019-02-15 Kristián Kozák , Bum Jun Kwon , Doowon Kim , Tudor Dumitraş

Authorship analysis is an important subject in the field of natural language processing. It allows the detection of the most likely writer of articles, news, books, or messages. This technique has multiple uses in tasks related to…

Pre-training, which utilizes extensive and varied datasets, is a critical factor in the success of Large Language Models (LLMs) across numerous applications. However, the detailed makeup of these datasets is often not disclosed, leading to…

Cryptography and Security · Computer Science 2024-01-02 Haodong Li , Gelei Deng , Yi Liu , Kailong Wang , Yuekang Li , Tianwei Zhang , Yang Liu , Guoai Xu , Guosheng Xu , Haoyu Wang

Recent advances in browser-based LLM agents have shown promise for automating tasks ranging from simple form filling to hotel booking or online shopping. Current benchmarks measure agent performance in controlled environments, such as…

Artificial Intelligence · Computer Science 2025-10-07 Su Kara , Fazle Faisal , Suman Nath

Large language model (LLM) agents are increasingly capable of autonomously conducting cyberattacks, posing significant threats to existing applications. This growing risk highlights the urgent need for a real-world benchmark to evaluate the…

Frauds severely hurt many kinds of Internet businesses. Group-based fraud detection is a popular methodology to catch fraudsters who unavoidably exhibit synchronized behaviors. We combine both graph-based features (e.g. cluster density) and…

Cryptography and Security · Computer Science 2018-06-26 Yikun Ban , Xin Liu , Tianyi Zhang , Ling Huang , Yitao Duan , Xue Liu , Wei Xu

Large Language Models (LLMs) frequently generate hallucinated content, posing significant challenges for applications where factuality is crucial. While existing hallucination detection methods typically operate at the sentence level or…

Machine Learning · Computer Science 2026-02-02 Albert Sawczyn , Jakub Binkowski , Denis Janiak , Bogdan Gabrys , Tomasz Kajdanowicz