English
Related papers

Related papers: What Large Language Models Do Not Talk About: An E…

200 papers

Large Language Models (LLMs) are trained to refuse to respond to harmful content. However, systematic analyses of whether this behavior is truly a reflection of its safety policies or an indication of political censorship, that is practiced…

Computation and Language · Computer Science 2025-12-01 Neemesh Yadav , Francesco Ortu , Jiarui Liu , Joeun Yook , Bernhard Schölkopf , Rada Mihalcea , Alberto Cazzaniga , Zhijing Jin

Large language models (LLMs) are trained on vast amounts of data to generate natural language, enabling them to perform tasks like text summarization and question answering. These models have become popular in artificial intelligence (AI)…

Large Language Models (LLMs) are a transformational technology, fundamentally changing how people obtain information and interact with the world. As people become increasingly reliant on them for an enormous variety of tasks, a body of…

Computers and Society · Computer Science 2025-05-08 Nouar Aldahoul , Hazem Ibrahim , Matteo Varvello , Aaron Kaufman , Talal Rahwan , Yasir Zaki

Large language models (LLMs) have transformed the way we access information. These models are often tuned to refuse to comply with requests that are considered harmful and to produce responses that better align with the preferences of those…

Computation and Language · Computer Science 2025-08-12 Hannah Cyberey , David Evans

Large language models (LLMs) have rapidly become indispensable tools for acquiring information and supporting human decision-making. However, ensuring that these models uphold fairness across varied contexts is critical to their safe and…

Computers and Society · Computer Science 2026-03-05 Xulang Zhang , Rui Mao , Erik Cambria

As internet access expands, so does exposure to harmful content, increasing the need for effective moderation. Research has demonstrated that large language models (LLMs) can be effectively utilized for social media moderation tasks,…

Computation and Language · Computer Science 2026-02-06 Hsuan-Yu Chou , Wajiha Naveed , Shuyan Zhou , Xiaowei Yang

Large language models (LLMs) have exhibited impressive capabilities in comprehending complex instructions. However, their blind adherence to provided instructions has led to concerns regarding risks of malicious use. Existing defence…

Artificial Intelligence · Computer Science 2023-07-25 David Glukhov , Ilia Shumailov , Yarin Gal , Nicolas Papernot , Vardan Papyan

Sensitive information detection is crucial in content moderation to maintain safe online communities. Assisting in this traditionally manual process could relieve human moderators from overwhelming and tedious tasks, allowing them to focus…

This study examines information suppression mechanisms in DeepSeek, an open-source large language model (LLM) developed in China. We propose an auditing framework and use it to analyze the model's responses to 646 politically sensitive…

Computers and Society · Computer Science 2025-06-17 Peiran Qiu , Siyi Zhou , Emilio Ferrara

Large language models' (LLMs') outputs are shaped by opaque and frequently-changing company content moderation policies and practices. LLM moderation often takes the form of refusal; models' refusal to produce text about certain topics both…

Computation and Language · Computer Science 2025-10-03 Yunlang Dai , Emma Lurie , Danaé Metaxa , Sorelle A. Friedler

We investigate the efficacy of Large Language Models (LLMs) in detecting implicit and explicit hate speech, examining how models with minimal safety alignment (uncensored) compare with more heavily aligned (censored) counterparts in a…

Computation and Language · Computer Science 2026-05-05 Sanjeeevan Selvaganapathy , Mehwish Nasim

DeepSeek recently released R1, a high-performing large language model (LLM) optimized for reasoning tasks. Despite its efficient training pipeline, R1 achieves competitive performance, even surpassing leading reasoning models like OpenAI's…

Computation and Language · Computer Science 2025-05-20 Ali Naseh , Harsh Chaudhari , Jaechul Roh , Mingshi Wu , Alina Oprea , Amir Houmansadr

As Large Language Models (LLMs) increasingly mediate global information access with the potential to shape public discourse, their alignment with universal human rights principles becomes important to ensure that these rights are abided by…

Computation and Language · Computer Science 2026-03-05 Keenan Samway , Nicole Miu Takagi , Rada Mihalcea , Bernhard Schölkopf , Ilias Chalkidis , Daniel Hershcovich , Zhijing Jin

Large Language Models (LLMs) have gained significant popularity for their application in various everyday tasks such as text generation, summarization, and information retrieval. As the widespread adoption of LLMs continues to surge, it…

Computation and Language · Computer Science 2024-03-22 Pagnarasmey Pit , Xingjun Ma , Mike Conway , Qingyu Chen , James Bailey , Henry Pit , Putrasmey Keo , Watey Diep , Yu-Gang Jiang

Large language models sometimes produce false or misleading responses. Two approaches to this problem are honesty elicitation -- modifying prompts or weights so that the model answers truthfully -- and lie detection -- classifying whether a…

Machine Learning · Computer Science 2026-03-11 Helena Casademunt , Bartosz Cywiński , Khoi Tran , Arya Jakkli , Samuel Marks , Neel Nanda

Large Language Models (LLMs) have emerged as powerful tools for generating human-like text, transforming human-machine interactions. However, their widespread adoption has raised concerns about their potential to influence public opinion…

Computers and Society · Computer Science 2025-03-24 Andre G. C. Pacheco , Athus Cavalini , Giovanni Comarela

The widespread dissemination of hate speech, harassment, harmful and sexual content, and violence across websites and media platforms presents substantial challenges and provokes widespread concern among different sectors of society.…

Computer Vision and Pattern Recognition · Computer Science 2024-11-27 Nouar AlDahoul , Myles Joshua Toledo Tan , Harishwar Reddy Kasireddy , Yasir Zaki

The ability of Natural Language Processing (NLP) methods to categorize text into multiple classes has motivated their use in online content moderation tasks, such as hate speech and fake news detection. However, there is limited…

Computation and Language · Computer Science 2025-03-11 Neemesh Yadav , Jiarui Liu , Francesco Ortu , Roya Ensafi , Zhijing Jin , Rada Mihalcea

Large Language Models (LLMs) have revolutionized artificial intelligence, demonstrating remarkable computational power and linguistic capabilities. However, these models are inherently prone to various biases stemming from their training…

Computation and Language · Computer Science 2025-02-14 Riccardo Cantini , Giada Cosenza , Alessio Orsino , Domenico Talia

Large language models (LLMs) are increasingly being utilised across a range of tasks and domains, with a burgeoning interest in their application within the field of journalism. This trend raises concerns due to our limited understanding of…

Computation and Language · Computer Science 2024-06-18 Filip Trhlik , Pontus Stenetorp
‹ Prev 1 2 3 10 Next ›