中文
相关论文

相关论文: Addressing Topic Leakage in Cross-Topic Evaluation…

200 篇论文

A central problem that has been researched for many years in the field of digital text forensics is the question whether two documents were written by the same author. Authorship verification (AV) is a research branch in this field that…

计算与语言 · 计算机科学 2020-07-09 Oren Halvani , Lukas Graner , Roey Regev

Authorship verification (AV) is a fundamental research task in digital text forensics, which addresses the problem of whether two texts were written by the same person. In recent years, a variety of AV methods have been proposed that focus…

计算与语言 · 计算机科学 2021-07-02 Oren Halvani , Lukas Graner

Authorship verification (AV) is a research subject in the field of digital text forensics that concerns itself with the question, whether two documents have been written by the same person. During the past two decades, an increasing number…

机器学习 · 计算机科学 2019-06-26 Oren Halvani , Christian Winter , Lukas Graner

The automatic verification of document authorships is important in various settings. Researchers are for example judged and compared by the amount and impact of their publications and public figures are confronted by their posts on social…

机器学习 · 计算机科学 2022-08-25 Maximilian Stubbemann , Gerd Stumme

A high degree of topical diversity is often considered to be an important characteristic of interesting text documents. A recent proposal for measuring topical diversity identifies three distributions for assessing the diversity of…

计算与语言 · 计算机科学 2018-10-15 Hosein Azarbonyad , Mostafa Dehghani , Tom Kenter , Maarten Marx , Jaap Kamps , Maarten de Rijke

Finding influential users in online social networks is an important problem with many possible useful applications. HITS and other link analysis methods, in particular, have been often used to identify hub and authority users in web graphs…

社会与信息网络 · 计算机科学 2018-02-21 Roy Ka-Wei Lee , Tuan-Anh Hoang , Ee-Peng Lim

We are addressing two fundamental problems in authorship verification (AV): Topic variability and miscalibration. Variations in the topic of two disputed texts are a major cause of error for most AV systems. In addition, it is observed that…

计算与语言 · 计算机科学 2021-06-22 Benedikt Boenninghoff , Dorothea Kolossa , Robert M. Nickel

The PAN 2020 authorship verification (AV) challenge focuses on a cross-topic/closed-set AV task over a collection of fanfiction texts. Fanfiction is a fan-written extension of a storyline in which a so-called fandom topic describes the…

计算与语言 · 计算机科学 2020-08-25 Benedikt Boenninghoff , Julian Rupp , Robert M. Nickel , Dorothea Kolossa

Authorship verification (AV) is the task of determining whether two texts were written by the same author and has been studied extensively, predominantly for English data. In contrast, large-scale benchmarks and systematic evaluations for…

计算与语言 · 计算机科学 2026-04-24 Lotta Kiefer , Christoph Leiter , Sotaro Takeshita , Elena Schmidt , Steffen Eger

Authorship attribution is the problem of identifying the most plausible author of an anonymous text from a set of candidate authors. Researchers have investigated same-topic and cross-topic scenarios of authorship attribution, which differ…

计算与语言 · 计算机科学 2021-09-10 Malik H. Altakrori , Jackie Chi Kit Cheung , Benjamin C. M. Fung

We adapt the Higher Criticism (HC) goodness-of-fit test to measure the closeness between word-frequency tables. We apply this measure to authorship attribution challenges, where the goal is to identify the author of a document using other…

计算与语言 · 计算机科学 2023-10-03 Alon Kipnis

Authorship attribution asks whether two pieces of text share a writer, but topical confound makes the task deceptively easy: two authors covering the same topic may look more alike than one author covering two topics. Scholarly prose offers…

数字图书馆 · 计算机科学 2026-05-26 Francis Kulumba , Wissam Antoun , Guillaume Vimont , Laurent Romary , Florian Cafiero

One of the main drivers of the recent advances in authorship verification is the PAN large-scale authorship dataset. Despite generating significant progress in the field, inconsistent performance differences between the closed and open test…

计算与语言 · 计算机科学 2022-11-02 Florin Brad , Andrei Manolache , Elena Burceanu , Antonio Barbalau , Radu Ionescu , Marius Popescu

Authorship Verification (AV) is a text classification task concerned with inferring whether a candidate text has been written by one specific author or by someone else. It has been shown that many AV systems are vulnerable to adversarial…

机器学习 · 计算机科学 2024-10-30 Silvia Corbara , Alejandro Moreo

Large Language Models (LLMs) are trained on massive web-crawled corpora. This poses risks of leakage, including personal information, copyrighted texts, and benchmark datasets. Such leakage leads to undermining human trust in AI due to…

计算与语言 · 计算机科学 2024-03-26 Masahiro Kaneko , Timothy Baldwin

The viral spread of fake news has caused great social harm, making fake news detection an urgent task. Current fake news detection methods rely heavily on text information by learning the extracted news content or writing style of internal…

社会与信息网络 · 计算机科学 2021-02-16 Yuxiang Ren , Jiawei Zhang

Linguistic style is an integral component of language. Recent advances in the development of style representations have increasingly used training objectives from authorship verification (AV): Do two texts have the same author? The…

计算与语言 · 计算机科学 2022-04-12 Anna Wegmann , Marijn Schraagen , Dong Nguyen

Authorship verification is the task of analyzing the linguistic patterns of two or more texts to determine whether they were written by the same author or not. The analysis is traditionally performed by experts who consider linguistic…

计算与语言 · 计算机科学 2019-11-21 Benedikt Boenninghoff , Steffen Hessler , Dorothea Kolossa , Robert M. Nickel

Authorship Verification (AV) (do two documents have the same author?) is essential in many real-life applications. AV is often used in privacy-sensitive domains that require an offline proprietary model that is deployed on premises, making…

计算与语言 · 计算机科学 2025-02-11 Sahana Ramnath , Kartik Pandey , Elizabeth Boschee , Xiang Ren

How much does a machine learning algorithm leak about its training data, and why? Membership inference attacks are used as an auditing tool to quantify this leakage. In this paper, we present a comprehensive \textit{hypothesis testing…

机器学习 · 计算机科学 2022-09-14 Jiayuan Ye , Aadyaa Maddi , Sasi Kumar Murakonda , Vincent Bindschaedler , Reza Shokri
‹ 上一页 1 2 3 10 下一页 ›