中文
相关论文

相关论文: Do Linear Probes Generalize Better in Persona Coor…

200 篇论文

Large language models (LLMs) have demonstrated remarkable capabilities in simulating human behaviour and social intelligence. However, they risk perpetuating societal biases, especially when demographic information is involved. We introduce…

计算机与社会 · 计算机科学 2025-06-11 Bryan Chen Zhengyu Tan , Roy Ka-Wei Lee

Pedestrian detection is used in many vision based applications ranging from video surveillance to autonomous driving. Despite achieving high performance, it is still largely unknown how well existing detectors generalize to unseen data.…

计算机视觉与模式识别 · 计算机科学 2020-12-10 Irtiza Hasan , Shengcai Liao , Jinpeng Li , Saad Ullah Akram , Ling Shao

In human-in-the-loop machine learning, the user provides information beyond that in the training data. Many algorithms and user interfaces have been designed to optimize and facilitate this human--machine interaction; however, fewer studies…

人机交互 · 计算机科学 2018-03-12 Pedram Daee , Tomi Peltola , Aki Vehtari , Samuel Kaski

Tracking the internal states of large language models across conversations is important for safety, interpretability, and model welfare, yet current methods are limited. Linear probes and other white-box methods compress high-dimensional…

人工智能 · 计算机科学 2026-04-14 Nicolas Martorell , Bruno Bianchi

Using offline datasets to evaluate conversational agents often fails to cover rare scenarios or to support testing new policies. This has motivated the use of controllable user simulators for targeted, counterfactual evaluation, typically…

人工智能 · 计算机科学 2026-05-13 Guy Tennenholtz , Ofer Meshi , Amir Globerson , Uri Shalit , Jihwan Jeong , Craig Boutilier

Behavioral logs provide rich signals for user modeling, but are noisy and interleaved across diverse intents. Recent work uses LLMs to generate interpretable natural-language personas from user logs, yet evaluation often emphasizes…

人工智能 · 计算机科学 2026-04-30 Nayoung Choi , Haeyu Jeong , Changbong Kim , Hongjun Lim , Jinho D. Choi

Large Language Models (LLMs) used in creative workflows can reinforce stereotypes and perpetuate inequities, making fairness auditing essential. Existing methods rely on constrained tasks and fixed benchmarks, leaving open-ended creative…

计算机与社会 · 计算机科学 2026-02-25 Hongliu Cao , Eoin Thomas , Rodrigo Acuna Agost

In this paper we generalize three identification recursive algorithms belonging to the pseudo-linear class, by introducing a predictor established on a generalized orthonormal function basis. Contrary to the existing identification schemes…

系统与控制 · 计算机科学 2019-08-14 Bernard Vau , Henri Bourlès

Use cases of sentiment analysis in the humanities often require contextualized, continuous scores. Concept Vector Projections (CVP) offer a recent solution: by modeling sentiment as a direction in embedding space, they produce continuous,…

计算与语言 · 计算机科学 2026-04-09 Laurits Lyngbaek , Pascale Feldkamp , Yuri Bizzoni , Kristoffer L. Nielbo , Kenneth Enevoldsen

When machine learning models are deployed on a test distribution different from the training distribution, they can perform poorly, but overestimate their performance. In this work, we aim to better estimate a model's performance under…

机器学习 · 计算机科学 2020-07-08 Ching-Yao Chuang , Antonio Torralba , Stefanie Jegelka

A personalized LLM should remember user facts, apply them correctly, and adapt over time to provide responses that the user prefers. Existing LLM personalization benchmarks are largely centered on two axes: accurately recalling user…

机器学习 · 计算机科学 2025-12-16 Md Awsafur Rahman , Adam Gabrys , Doug Kang , Jingjing Sun , Tian Tan , Ashwin Chandramouli

Sparse principal component analysis (sparse PCA) is a widely used technique for dimensionality reduction in multivariate analysis, addressing two key limitations of standard PCA. First, sparse PCA can be implemented in high-dimensional low…

统计方法学 · 统计学 2025-10-07 Jan O. Bauer

Given the claims of improved text generation quality across various pre-trained neural models, we consider the coherence evaluation of machine generated text to be one of the principal applications of coherence models that needs to be…

计算与语言 · 计算机科学 2022-03-22 Prathyusha Jwalapuram , Shafiq Joty , Xiang Lin

Some traits making a "good" AI model are hard to describe upfront. For example, should responses be more polite or more casual? Such traits are sometimes summarized as model character or personality. Without a clear objective, conventional…

计算与语言 · 计算机科学 2025-10-01 Arduin Findeis , Timo Kaufmann , Eyke Hüllermeier , Robert Mullins

Toxicity detection is inherently subjective, shaped by the diverse perspectives and social priors of different demographic groups. While ``pluralistic'' modeling as used in economics and the social sciences aims to capture perspective…

计算与语言 · 计算机科学 2026-01-06 Berk Atil , Rebecca J. Passonneau , Ninareh Mehrabi

Overparameterization in deep learning is powerful: Very large models fit the training data perfectly and yet often generalize well. This realization brought back the study of linear models for regression, including ordinary least squares…

机器学习 · 统计学 2022-04-07 Ningyuan Huang , David W. Hogg , Soledad Villar

Driven by the demand for personalized AI systems, there is growing interest in aligning the behavior of large language models (LLMs) with human traits such as personality. Previous attempts to induce personality in LLMs have shown promising…

计算与语言 · 计算机科学 2025-09-25 Seungjong Sun , Seo Yeon Baek , Jang Hyun Kim

A common method to study deep learning systems is to use simplified model representations--for example, using singular value decomposition to visualize the model's hidden states in a lower dimensional space. This approach assumes that the…

机器学习 · 计算机科学 2024-06-06 Dan Friedman , Andrew Lampinen , Lucas Dixon , Danqi Chen , Asma Ghandeharioun

Do large language models (LLMs) anticipate when they will answer correctly? To study this, we extract activations after a question is read but before any tokens are generated, and train linear probes to predict whether the model's…

The influence of personas on Large Language Models (LLMs) has been widely studied, yet their direct impact on performance remains uncertain. This work explores a novel approach to guiding LLM behaviour through role vectors, an alternative…

计算与语言 · 计算机科学 2025-02-18 Daniele Potertì , Andrea Seveso , Fabio Mercorio