中文
相关论文

相关论文: The Topic Confusion Task: A Novel Scenario for Aut…

200 篇论文

The Chaos Game Representation, a method for creating images from nucleotide sequences, is modified to make images from chunks of text documents. Machine learning methods are then applied to train classifiers based on authorship. Experiments…

计算与语言 · 计算机科学 2018-02-19 Daniel Lichtblau , Catalin Stoean

Text-based analysis methods allow to reveal privacy relevant author attributes such as gender, age and identify of the text's author. Such methods can compromise the privacy of an anonymous author even when the author tries to remove…

密码学与安全 · 计算机科学 2018-02-20 Rakshith Shetty , Bernt Schiele , Mario Fritz

Mis/disinformation is a common and dangerous occurrence on social media. Misattribution is a form of mis/disinformation that deals with a false claim of authorship, which means a user is claiming someone said (posted) something they never…

信息检索 · 计算机科学 2024-10-10 Ashlyn M. Farris , Michael L. Nelson

This paper elaborates on the notion of uncertainty in the context of annotation in large text corpora, specifically focusing on (but not limited to) historical languages. Such uncertainty might be due to inherent properties of the language,…

计算与语言 · 计算机科学 2021-05-31 Marie-Luis Merten , Marcel Wever , Michaela Geierhos , Doris Tophinke , Eyke Hüllermeier

Most of privacy protection studies for textual data focus on removing explicit sensitive identifiers. However, personal writing style, as a strong indicator of the authorship, is often neglected. Recent studies, such as SynTF, have shown…

密码学与安全 · 计算机科学 2021-05-14 Haohan Bo , Steven H. H. Ding , Benjamin C. M. Fung , Farkhund Iqbal

Paraphrase Identification is a fundamental task in Natural Language Processing. While much progress has been made in the field, the performance of many state-of-the-art models often suffer from distribution shift during inference time. We…

计算与语言 · 计算机科学 2022-10-06 Yifei Zhou , Renyu Li , Hayden Housen , Ser-Nam Lim

Text-embedding models often exhibit biases arising from the data on which they are trained. In this paper, we examine a hitherto unexplored bias in text-embeddings: bias arising from the presence of $\textit{names}$ such as persons,…

计算与语言 · 计算机科学 2025-02-06 Sahil Manchanda , Pannaga Shivaswamy

Recent advancements in language representation models such as BERT have led to a rapid improvement in numerous natural language processing tasks. However, language models usually consist of a few hundred million trainable parameters with…

机器学习 · 计算机科学 2019-12-12 Mehrdad Valipour , En-Shiun Annie Lee , Jaime R. Jamacaro , Carolina Bessega

Over the years there has been ongoing interest in detecting authorship of a text based on statistical properties of the text, such as by using occurrence rates of noncontextual words. In previous work, these techniques have been used, for…

计算与语言 · 计算机科学 2024-03-21 Todd K Moon , Jacob H. Gunther

Assigning qualified, unbiased and interested reviewers to paper submissions is vital for maintaining the integrity and quality of the academic publishing system and providing valuable reviews to authors. However, matching thousands of…

信息检索 · 计算机科学 2022-11-09 Omer Anjum , Alok Kamatar , Toby Liang , Jinjun Xiong , Wen-mei Hwu

Modern generative search engines enhance the reliability of large language model (LLM) responses by providing cited evidence. However, evaluating the answer's attribution, i.e., whether every claim within the generated responses is fully…

计算与语言 · 计算机科学 2024-02-26 Yifei Li , Xiang Yue , Zeyi Liao , Huan Sun

Recently, topic modeling has been widely used to discover the abstract topics in text corpora. Most of the existing topic models are based on the assumption of three-layer hierarchical Bayesian structure, i.e. each document is modeled as a…

计算与语言 · 计算机科学 2017-04-10 Yi-Kun Tang , Xian-Ling Mao , Heyan Huang , Guihua Wen

Large Language Models (LLMs), such as GPT-4 and Llama, have demonstrated remarkable abilities in generating natural language. However, they also pose security and integrity challenges. Existing countermeasures primarily focus on…

密码学与安全 · 计算机科学 2025-08-21 Zixin Rao , Youssef Mohamed , Shang Liu , Zeyan Liu

We propose a formal definition for the task of suggestion mining in the context of a wide range of open domain applications. Human perception of the term \emph{suggestion} is subjective and this effects the preparation of hand labeled…

计算与语言 · 计算机科学 2018-07-03 Sapna Negi , Maarten de Rijke , Paul Buitelaar

Transfer learning methods, and in particular domain adaptation, help exploit labeled data in one domain to improve the performance of a certain task in another domain. However, it is still not clear what factors affect the success of domain…

计算与语言 · 计算机科学 2021-06-25 Nicolai Pogrebnyakov , Shohreh Shaghaghian

Text is a vehicle to convey information that reflects the writer's linguistic style and communicative patterns. By studying these attributes, we can discover latent insights about the author and their underlying message. This article uses…

计算机与社会 · 计算机科学 2024-12-19 Deborah Gerhardt , Miriam Marcowitz-Bitton , W. Michael Schuster , Avshalom Elmalech , Omri Suissa , Moshe Mash

Authorship identification ascertains the authorship of texts whose origins remain undisclosed. That authorship identification techniques work as reliably as they do has been attributed to the fact that authorial style is properly captured…

计算与语言 · 计算机科学 2023-10-03 Haining Wang

Leveraging short-text contents to estimate the occupation of microblog authors has significant gains in many applications. Yet challenges abound. Firstly brief textual contents come with excessive lexical noise that makes the inference…

计算与语言 · 计算机科学 2021-06-15 Sayna Esmailzadeh , Saeid Hosseini , Mohammad Reza Kangavari , Wen Hua

Uncertainty Quantification (UQ) research has primarily focused on closed-book factual question answering (QA), while contextual QA remains unexplored, despite its importance in real-world applications. In this work, we focus on UQ for the…

Unsupervised text attribute transfer automatically transforms a text to alter a specific attribute (e.g. sentiment) without using any parallel data, while simultaneously preserving its attribute-independent content. The dominant approaches…

计算与语言 · 计算机科学 2019-12-13 Ke Wang , Hang Hua , Xiaojun Wan
‹ 上一页 1 8 9 10 下一页 ›