English
Related papers

Related papers: Mitigating Covertly Unsafe Text within Natural Lan…

200 papers

Recent research points toward LLMs being manipulated through adversarial and seemingly benign inputs, resulting in harmful, biased, or policy-violating outputs. In this paper, we study an underexplored issue concerning harmful and toxic…

Computation and Language · Computer Science 2026-03-27 Sagnik Basu , Subhrajit Mitra , Aman Juneja , Somnath Banerjee , Rima Hazra , Animesh Mukherjee

Technological advancements have resulted in an exponential increase in the use of online social networks (OSNs) worldwide. While online social networks provide a great communication medium, they also increase the user's exposure to…

Computers and Society · Computer Science 2024-01-09 Sylvia W Azumah , Nelly Elsayed , Zag ElSayed , Murat Ozer

The exploding rate of data publishing in our networked society has magnified the risk of sensitive information leakage and misuse, pushing the need to secure multimedia content from unintended exposure to potentially untrusted third…

Cryptography and Security · Computer Science 2025-09-16 Andrea Ciccotelli , Hanaa Abbas , Roberto Di Pietro

Language models trained on large-scale unfiltered datasets curated from the open web acquire systemic biases, prejudices, and harmful views from their training data. We present a methodology for programmatically identifying and removing…

Computation and Language · Computer Science 2021-11-30 Helen Ngo , Cooper Raterink , João G. M. Araújo , Ivan Zhang , Carol Chen , Adrien Morisot , Nicholas Frosst

This work examines how leading generative artificial intelligence companies construct and communicate the concept of "safety" through public-facing documents. Drawing on critical discourse analysis, we analyze a corpus of corporate…

Computers and Society · Computer Science 2026-02-10 Ankolika De , Gabriel Lima , Yixin Zou

Models trained on large unlabeled corpora of human interactions will learn patterns and mimic behaviors therein, which include offensive or otherwise toxic behavior and unwanted biases. We investigate a variety of methods to mitigate these…

Computation and Language · Computer Science 2021-08-06 Jing Xu , Da Ju , Margaret Li , Y-Lan Boureau , Jason Weston , Emily Dinan

The recent surge in high-quality open-source Generative AI text models (colloquially: LLMs), as well as efficient finetuning techniques, have opened the possibility of creating high-quality personalized models that generate text attuned to…

Computation and Language · Computer Science 2025-06-03 Eugenia Iofinova , Andrej Jovanovic , Dan Alistarh

Nowadays, the spread of misinformation is a prominent problem in society. Our research focuses on aiding the automatic identification of misinformation by analyzing the persuasive strategies employed in textual documents. We introduce a…

Computation and Language · Computer Science 2024-04-11 Danial Kamali , Joseph Romain , Huiyi Liu , Wei Peng , Jingbo Meng , Parisa Kordjamshidi

As Large Language Models (LLMs) continue to advance in understanding and generating long sequences, new safety concerns have been introduced through the long context. However, the safety of LLMs in long-context tasks remains under-explored,…

Computation and Language · Computer Science 2025-02-25 Yida Lu , Jiale Cheng , Zhexin Zhang , Shiyao Cui , Cunxiang Wang , Xiaotao Gu , Yuxiao Dong , Jie Tang , Hongning Wang , Minlie Huang

Automated simplification models aim to make input texts more readable. Such methods have the potential to make complex information accessible to a wider audience, e.g., providing access to recent medical literature which might otherwise be…

Computation and Language · Computer Science 2022-04-18 Ashwin Devaraj , William Sheffield , Byron C. Wallace , Junyi Jessy Li

Textual deception constitutes a major problem for online security. Many studies have argued that deceptiveness leaves traces in writing style, which could be detected using text classification techniques. By conducting an extensive…

Computation and Language · Computer Science 2019-02-27 Tommi Gröndahl , N. Asokan

Large language models (LLMs) have emerged as an integral part of modern societies, powering user-facing applications such as personal assistants and enterprise applications like recruitment tools. Despite their utility, research indicates…

Computation and Language · Computer Science 2024-05-10 Preetam Prabhu Srikar Dammu , Hayoung Jung , Anjali Singh , Monojit Choudhury , Tanushree Mitra

The advent of Large Language Models (LLMs) has made a transformative impact. However, the potential that LLMs such as ChatGPT can be exploited to generate misinformation has posed a serious concern to online safety and public trust. A…

Computation and Language · Computer Science 2024-04-25 Canyu Chen , Kai Shu

Large language models (LLMs) have advanced to a point that even humans have difficulty discerning whether a text was generated by another human, or by a computer. However, knowing whether a text was produced by human or artificial…

Computation and Language · Computer Science 2025-04-15 Kathleen C. Fraser , Hillary Dawkins , Svetlana Kiritchenko

The rapid advancement in neurotechnology in recent years has created an emerging critical intersection between neurotechnology and security. Implantable devices, non-invasive monitoring, and non-invasive therapies all carry with them the…

Cryptography and Security · Computer Science 2025-01-28 Bryce Allen Bagley , Claudia K Petritsch

Spurred by the recent rapid increase in the development and distribution of large language models (LLMs) across industry and academia, much recent work has drawn attention to safety- and security-related threats and vulnerabilities of LLMs,…

Computation and Language · Computer Science 2023-08-25 Maximilian Mozes , Xuanli He , Bennett Kleinberg , Lewis D. Griffin

Large language models exhibit safety degradation in non-English languages. Standard evaluation relies on Jailbreak Success Rate (JSR), which confounds several safety-driving factors into one, obscuring the specific cause(s) of safety…

Computation and Language · Computer Science 2026-05-19 Max Zhang , Ameen Patel , Sang T. Truong , Sanmi Koyejo

With the development of large language models (LLMs), detecting whether text is generated by a machine becomes increasingly challenging in the face of malicious use cases like the spread of false information, protection of intellectual…

Computation and Language · Computer Science 2024-04-03 Ying Zhou , Ben He , Le Sun

Probabilistic text generators have been used to produce fake scientific papers for more than a decade. Such nonsensical papers are easily detected by both human and machine. Now more complex AI-powered generation techniques produce texts…

Digital Libraries · Computer Science 2021-07-15 Guillaume Cabanac , Cyril Labbé , Alexander Magazinov

As large language models become more prevalent, their possible harmful or inappropriate responses are a cause for concern. This paper introduces a unique dataset containing adversarial examples in the form of questions, which we call AttaQ,…

Computation and Language · Computer Science 2023-11-08 George Kour , Marcel Zalmanovici , Naama Zwerdling , Esther Goldbraich , Ora Nova Fandina , Ateret Anaby-Tavor , Orna Raz , Eitan Farchi