中文
相关论文

相关论文: Depth-Wise Activation Steering for Honest Language…

200 篇论文

Large Language Models (LLMs) are widely deployed in reasoning, planning, and decision-making tasks, making their trustworthiness critical. A significant and underexplored risk is intentional deception, where an LLM deliberately fabricates…

机器学习 · 计算机科学 2026-05-04 Zhaomin Wu , Mingzhe Du , See-Kiong Ng , Bingsheng He

Activation steering provides parameter-efficient control over large language models (LLMs) at inference time, but many methods rely on off-distribution supervision and discrete masking, leading to brittle interventions. We propose ROAST…

机器学习 · 计算机科学 2026-02-17 Xuanbo Su , Hao Luo , Yingfang Zhang , Lijun Zhang

Large Language Models (LLMs) have revolutionised natural language processing, exhibiting impressive human-like capabilities. In particular, LLMs are capable of "lying", knowingly outputting false statements. Hence, it is of interest and…

计算与语言 · 计算机科学 2024-10-22 Lennart Bürger , Fred A. Hamprecht , Boaz Nadler

As autonomous agents become more capable of performing real-world tasks, distinguishing scheming behavior from benign task pursuit may become a central AI control problem. Existing monitors often rely on chain-of-thought access or internal…

Diffusion large language models (dLLMs) generate text via iterative denoising but consistently underperform on multi-step reasoning. We hypothesize this gap stems from a coordination problem: AR models build coherence token-by-token, while…

人工智能 · 计算机科学 2026-03-17 Earl J St Sauver

Multimodal large language models can exhibit text dominance, over-relying on linguistic priors instead of grounding predictions in non-text inputs. One example is large audio-language models (LALMs) where decisive audio evidence can be…

声音 · 计算机科学 2026-03-10 Neta Glazer , Lenny Aharon , Ethan Fetaya

Recent studies on post-training large language models (LLMs) for reasoning through reinforcement learning (RL) typically focus on tasks that can be accurately verified and rewarded, such as solving math problems. In contrast, our research…

计算与语言 · 计算机科学 2025-05-29 Ang Lv , Ruobing Xie , Xingwu Sun , Zhanhui Kang , Rui Yan

Large Language Models (LLMs) are trained on diverse and often conflicting knowledge spanning multiple domains and time periods. Some of this knowledge is only valid within specific temporal contexts, such as answering the question, "Who is…

计算与语言 · 计算机科学 2025-11-11 Sanjay Govindan , Maurice Pagnucco , Yang Song

Recent advances in Large Language Models (LLMs) have demonstrated remarkable progress in their reasoning capabilities, such as Chain-of-Thought (CoT). Most approaches rely on CoT rationales. Previous studies have shown that LLMs often…

计算与语言 · 计算机科学 2026-01-21 Kentaro Kazama , Daiki Shirafuji , Tatsuhiko Saito

Despite rapid progress in large language models (LLMs), the statistical structure of their weights, activations, and gradients-and its implications for initialization, training dynamics, and efficiency-remains largely unexplored. We…

机器学习 · 计算机科学 2026-02-24 Jun Wu , Patrick Huang , Jiangtao Wen , Yuxing Han

Research into 6G networks has been initiated to support a variety of critical artificial intelligence (AI) assisted applications such as autonomous driving. In such applications, AI-based decisions should be performed in a real-time manner.…

人工智能 · 计算机科学 2023-12-07 Abdul Karim Gizzini , Yahia Medjahdi , Ali J. Ghandour , Laurent Clavier

Activation steering methods modify intermediate representations of language models to control output behavior, but universally assume the activation space is Euclidean. We show this assumption fails drastically: the local geometry induced…

机器学习 · 计算机科学 2026-05-19 Sihan Wang , Jiayi Zhao

In statistical dialogue management, the dialogue manager learns a policy that maps a belief state to an action for the system to perform. Efficient exploration is key to successful policy optimisation. Current deep reinforcement learning…

机器学习 · 统计学 2017-12-04 Christopher Tegho , Paweł Budzianowski , Milica Gašić

Common methods for aligning large language models (LLMs) with desired behaviour heavily rely on human-labelled data. However, as models grow increasingly sophisticated, they will surpass human expertise, and the role of human evaluation…

Activation steering is a method for controlling Large Language Model (LLM) behavior by intervening in its internal representations to increase the alignment with a specific target feature direction. However, standard interventions, such as…

机器学习 · 计算机科学 2026-05-05 Tam Nguyen , Tu Anh Nguyen , Sina Alemohammad , Richard G. Baraniuk

In this study, we address the challenge of enabling large language models (LLMs) to consistently adhere to emotional support strategies in extended conversations. We focus on the steerability of the Llama-2 and Llama-3 suite of models,…

计算与语言 · 计算机科学 2024-09-17 Navid Madani , Sougata Saha , Rohini Srihari

As large language models (LLMs) evolve in complexity and capability, the efficacy of less widely deployed alignment techniques are uncertain. Building on previous work on activation steering and contrastive activation addition (CAA), this…

机器学习 · 计算机科学 2025-07-17 Sheikh Abdur Raheem Ali , Justin Xu , Ivory Yang , Jasmine Xinze Li , Ayse Arslan , Clark Benham

Large language models (LLMs) tend to verbalize confidence scores that are largely detached from their actual accuracy, yet the geometric relationship governing this behavior remain poorly understood. In this work, we present a mechanistic…

计算与语言 · 计算机科学 2026-04-02 Miranda Muqing Miao , Lyle Ungar

The use of large language models (LLMs) is expanding rapidly, and open-source versions are becoming available, offering users safer and more adaptable options. These models enable users to protect data privacy by eliminating the need to…

机器学习 · 计算机科学 2024-08-06 Hui Yin , Amir Aryani , Nakul Nambiar

Activation-based linear probing is widely proposed as a method for both detecting and correcting hallucinations in autoregressive language models. We present an empirical study across seven models spanning 117M to 7B parameters and three…

计算与语言 · 计算机科学 2026-05-12 Dip Roy , Rajiv Misra , Sanjay Kumar Singh , Anisha Roy