English
Related papers

Related papers: Do Linear Probes Generalize Better in Persona Coor…

200 papers

Reasoning in humans is prone to biases due to underlying motivations like identity protection, that undermine rational decision-making and judgment. This \textit{motivated reasoning} at a collective level can be detrimental to society when…

Artificial Intelligence · Computer Science 2026-04-20 Saloni Dash , Amélie Reymond , Emma S. Spiro , Aylin Caliskan

Language Models (LMs) may acquire harmful knowledge, and yet feign ignorance of these topics when under audit. Inspired by the recent discovery of deception-related behaviour patterns in LMs, we aim to train classifiers that detect when a…

Computation and Language · Computer Science 2026-03-24 Dhananjay Ashok , Ruth-Ann Armstrong , Jonathan May

Pre-trained language models (PLMs) are trained on data that inherently contains gender biases, leading to undesirable impacts. Traditional debiasing methods often rely on external corpora, which may lack quality, diversity, or demographic…

Computation and Language · Computer Science 2025-03-13 Liu Yu , Ludie Guo , Ping Kuang , Fan Zhou

Many text corpora exhibit socially problematic biases, which can be propagated or amplified in the models trained on such data. For example, doctor cooccurs more frequently with male pronouns than female pronouns. In this study we (i)…

Computation and Language · Computer Science 2019-04-08 Shikha Bordia , Samuel R. Bowman

Humans readily generalize, applying prior knowledge to novel situations and stimuli. Advances in machine learning and artificial intelligence have begun to approximate and even surpass human performance, but machine systems reliably…

Artificial Intelligence · Computer Science 2025-12-10 Leonidas A. A. Doumas , Guillermo Puebla , Andrea E. Martin

Poor generalization is one symptom of models that learn to predict target variables using spuriously-correlated image features present only in the training distribution instead of the true image features that denote a class. It is often…

Computer Vision and Pattern Recognition · Computer Science 2021-02-11 Joseph D. Viviano , Becks Simpson , Francis Dutil , Yoshua Bengio , Joseph Paul Cohen

Warning: This research studies AI persuasion and bias amplification that could be misused; all experiments are for safety evaluation. Large Language Models (LLMs) now generate convincing, human-like text and are widely used in content…

Computation and Language · Computer Science 2025-08-25 Saumya Roy

Large Language Models (LLMs) are increasingly deployed in interactions where they are prompted to adopt personas. This paper investigates whether such persona conditioning affects model safety under bullying, an adversarial manipulation…

Artificial Intelligence · Computer Science 2025-05-20 Ziwei Xu , Udit Sanghi , Mohan Kankanhalli

Using prompted language models as classifiers enables classification in domains with limited training data, but misses some of the robustness and performance benefits that fine-tuning can bring. We study whether training on multiple…

Artificial Intelligence · Computer Science 2026-05-13 Sam Martin , Fabien Roger

Several recent works argue that LLMs have a universal truth direction where true and false statements are linearly separable in the activation space of the model. It has been demonstrated that linear probes trained on a single hidden state…

Computation and Language · Computer Science 2025-05-16 Timour Ichmoukhamedov , David Martens

While advances in fairness and alignment have helped mitigate overt biases exhibited by large language models (LLMs) when explicitly prompted, we hypothesize that these models may still exhibit implicit biases when simulating human…

Computation and Language · Computer Science 2025-01-30 Yuxuan Li , Hirokazu Shirado , Sauvik Das

Weak-to-Strong Generalization (W2SG), where a weak model supervises a stronger one, serves as an important analogy for understanding how humans might guide superhuman intelligence in the future. Promising empirical results revealed that a…

Machine Learning · Computer Science 2025-06-19 Yihao Xue , Jiping Li , Baharan Mirzasoleiman

Neural representations are not unique objects. Even when two systems realize the same downstream computation, their hidden coordinates may differ by reparameterization. A probe family intended to reveal structure already present in a…

Machine Learning · Computer Science 2026-05-13 Su Hyeong Lee , Risi Kondor

A prominent issue in aligning language models (LMs) to personalized preferences is underspecification -- the lack of information from users about their preferences. A popular trend of injecting such specification is adding a prefix (e.g.…

Computation and Language · Computer Science 2025-09-30 Zilu Tang , Afra Feyza Akyürek , Ekin Akyürek , Derry Wijaya

Aging and chronic conditions affect older adults' daily lives, making the early detection of developing health issues crucial. Weakness, which is common across many conditions, can subtly alter physical movements and daily activities.…

Computer Vision and Pattern Recognition · Computer Science 2026-03-31 Chen Long-fei , Muhammad Ahmed Raza , Craig Innes , Subramanian Ramamoorthy , Robert B. Fisher

Detecting AI-generated images (AIGI) remains challenging because detectors often fail to generalize to unseen generators. Although existing methods are trained on large datasets, their performance still degrades when generation settings…

Computer Vision and Pattern Recognition · Computer Science 2026-05-28 Zijie Cao , Weijie Tu , Yao Xiao , Weijian Deng , Liang Lin , Pengxu Wei

Large NLP models have recently shown impressive performance in language understanding tasks, typically evaluated by their fine-tuned performance. Alternatively, probing has received increasing attention as being a lightweight method for…

Computation and Language · Computer Science 2022-10-17 Zining Zhu , Soroosh Shahtalebi , Frank Rudzicz

Neural language models (LMs) are arguably less data-efficient than humans from a language acquisition perspective. One fundamental question is why this human-LM gap arises. This study explores the advantage of grounded language acquisition,…

Computation and Language · Computer Science 2024-12-18 Tatsuki Kuribayashi , Timothy Baldwin

Activation-based steering can personalize large language models at inference time, but its effects in educational settings remain unclear. We study persona vectors for seven character traits in short-answer generation and automated scoring…

Computation and Language · Computer Science 2026-04-09 Yongchao Wu , Aron Henriksson

Human behavior models are essential as behavior references and for simulating human agents in virtual safety assessment of automated vehicles (AVs), yet current models face a trade-off between interpretability and flexibility.…

Artificial Intelligence · Computer Science 2026-05-19 Samir H. A. Mohammad , Wouter Mooi , Arkady Zgonnikov
‹ Prev 1 4 5 6 7 8 10 Next ›