English
Related papers

Related papers: Understanding and Mitigating Political Stance Cros…

200 papers

Large Language Models (LLMs) are increasingly deployed as gateways to information, yet their content moderation practices remain underexplored. This work investigates the extent to which LLMs refuse to answer or omit information when…

Computation and Language · Computer Science 2025-04-08 Sander Noels , Guillaume Bied , Maarten Buyl , Alexander Rogiers , Yousra Fettach , Jefrey Lijffijt , Tijl De Bie

Instruction-following language models often show undesirable biases. These undesirable biases may be accelerated in the real-world usage of language models, where a wide range of instructions is used through zero-shot example prompting. To…

Artificial Intelligence · Computer Science 2024-06-06 Nakyeong Yang , Taegwan Kang , Jungkyu Choi , Honglak Lee , Kyomin Jung

Multilingual large language models (LLMs) achieve strong performance across languages, yet how language capabilities are organized at the neuron level remains poorly understood. Prior work has identified language-related neurons mainly…

Computation and Language · Computer Science 2026-03-11 Yifan Le , Yunliang Li

How to usefully encode compositional task structure has long been a core challenge in AI. Recent work in chain of thought prompting has shown that for very large neural language models (LMs), explicitly demonstrating the inferential steps…

Computation and Language · Computer Science 2022-10-25 Victor S. Bursztyn , David Demeter , Doug Downey , Larry Birnbaum

Large language models (LLMs) are increasingly deployed in politically sensitive settings, raising concerns about their potential to encode, amplify, or be steered toward specific ideologies. We investigate how adopting synthetic personas…

Computation and Language · Computer Science 2025-08-25 Pietro Bernardelle , Stefano Civelli , Leon Fröhling , Riccardo Lunardi , Kevin Roitero , Gianluca Demartini

Large language models (LLMs) are increasingly used in everyday tools and applications, raising concerns about their potential influence on political views. While prior research has shown that LLMs often exhibit measurable political…

Computation and Language · Computer Science 2025-11-03 Daniil Gurgurov , Katharina Trinley , Ivan Vykopal , Josef van Genabith , Simon Ostermann , Roberto Zamparelli

Large Language Models (LLMs) have demonstrated remarkable capabilities across a wide range of instruction-following tasks, yet their grasp of nuanced social science concepts remains underexplored. This paper examines whether LLMs can…

Computation and Language · Computer Science 2025-07-28 Ilias Chalkidis , Stephanie Brandl , Paris Aslanidis

We study how targeted content injection can strategically disrupt social networks. Using the Friedkin-Johnsen (FJ) model, we utilize a measure of social dissensus and show that (i) simple FJ variants cannot significantly perturb the…

Social and Information Networks · Computer Science 2025-11-03 Erica Coppolillo , Giuseppe Manco

Recent experiments revealed that a certain class of inhibitory neurons in the cerebral cortex make synapses not onto cell bodies but at distal parts of dendrites of the target neurons, mediating highly nonlinear dendritic inhibition. We…

Neurons and Cognition · Quantitative Biology 2007-05-23 Kenji Morita , Kazuyuki Aihara

LVLMs achieve remarkable multimodal understanding and generation but remain susceptible to hallucinations. Existing mitigation methods predominantly focus on output-level adjustments, leaving the internal mechanisms that give rise to these…

Computer Vision and Pattern Recognition · Computer Science 2026-02-03 Guangtao Lyu , Xinyi Cheng , Qi Liu , Chenghao Xu , Jiexi Yan , Muli Yang , Fen Fang , Cheng Deng

Developing machine learning models to characterize political polarization on online social media presents significant challenges. These challenges mainly stem from various factors such as the lack of annotated data, presence of noise in…

Social and Information Networks · Computer Science 2023-11-22 Sadia Kamal , Brenner Little , Jade Gullic , Trevor Harms , Kristin Olofsson , Arunkumar Bagavathi

Topic models and all their variants analyse text by learning meaningful representations through word co-occurrences. As pointed out by Williamson et al. (2010), such models implicitly assume that the probability of a topic to be active and…

Computation and Language · Computer Science 2023-01-27 Kostadin Cvejoski , Ramsés J. Sánchez , César Ojeda

This study investigates the adoption of open-access, locally deployable causal large language models (LLMs) for travel mode choice prediction and introduces LiTransMC, the first fine-tuned causal LLM developed for this task. We…

Computation and Language · Computer Science 2025-10-08 Tareq Alsaleh , Bilal Farooq

Reinforcement learning (RL) has become a key technique for enhancing the reasoning abilities of large language models (LLMs), with policy-gradient algorithms dominating the post-training stage because of their efficiency and effectiveness.…

Artificial Intelligence · Computer Science 2025-08-08 Chang Tian , Matthew B. Blaschko , Mingzhe Xing , Xiuxing Li , Yinliang Yue , Marie-Francine Moens

Supervised Fine-Tuning (SFT) and Reinforcement Learning (RL) are two prominent post-training paradigms for refining the capabilities and aligning the behavior of Large Language Models (LLMs). Existing approaches that integrate SFT and RL…

Machine Learning · Computer Science 2026-03-18 Wenhao Zhang , Yuexiang Xie , Yuchang Sun , Yanxi Chen , Guoyin Wang , Yaliang Li , Bolin Ding , Jingren Zhou

Fine-tuning large language models (LLMs) can lead to unintended out-of-distribution generalization. Standard approaches to this problem rely on modifying training data, for example by adding data that better specify the intended…

Machine Learning · Computer Science 2025-11-11 Helena Casademunt , Caden Juang , Adam Karvonen , Samuel Marks , Senthooran Rajamanoharan , Neel Nanda

During training, Large Language Models (LLMs) learn social regularities that can lead to gender bias in downstream applications. Most mitigation efforts focus on reducing bias in generated outputs, typically evaluated on structured…

Computation and Language · Computer Science 2026-05-14 Nour Bouchouchi , Thibault Laugel , Xavier Renard , Christophe Marsala , Marie-Jeanne Lesot , Marcin Detyniecki

Deep Neural Networks are prone to learning spurious correlations embedded in the training data, leading to potentially biased predictions. This poses risks when deploying these models for high-stake decision-making, such as in medical…

Machine Learning · Computer Science 2023-12-19 Maximilian Dreyer , Frederik Pahde , Christopher J. Anders , Wojciech Samek , Sebastian Lapuschkin

The training of large language models (LLMs) on extensive, unfiltered corpora sourced from the internet is a common and advantageous practice. Consequently, LLMs have learned and inadvertently reproduced various types of biases, including…

Computation and Language · Computer Science 2023-11-20 Ambri Ma , Arnav Kumar , Brett Zeligson

We localize the policy routing mechanism in alignment-trained language models. An intermediate-layer attention gate reads detected content and triggers deeper amplifier heads that boost the signal toward refusal. In smaller models the gate…

Computation and Language · Computer Science 2026-05-04 Gregory N. Frank