中文
相关论文

相关论文: Steering Awareness: Detecting Activation Steering …

200 篇论文

This paper presents a novel approach to Autonomous Vehicle (AV) control through the application of active inference, a theory derived from neuroscience that conceptualizes the brain as a predictive machine. Traditional autonomous driving…

机器人学 · 计算机科学 2025-03-17 Elahe Delavari , John Moore , Junho Hong , Jaerock Kwon

Steering models (such as the generalized two-point model) predict human steering behavior well when the human is in direct control of a vehicle. In vehicles under autonomous control, human control inputs are not used; rather, an autonomous…

系统与控制 · 电气工程与系统科学 2024-10-02 Rene Mai , Agung Julius , Sandipan Mishra

Deep neural perception and control networks are likely to be a key component of self-driving vehicles. These models need to be explainable - they should provide easy-to-interpret rationales for their behavior - so that passengers, insurance…

计算机视觉与模式识别 · 计算机科学 2017-04-03 Jinkyu Kim , John Canny

Driver distraction is a well-known cause for traffic collisions worldwide. Studies have indicated that shared steering control, which actively provides haptic guidance torque on the steering wheel, effectively improves the performance of…

系统与控制 · 电气工程与系统科学 2021-06-08 Zheng Wang , Satoshi Suga , Edric John Cruz Nacpil , Bo Yang , Kimihiko Nakano

Reliable behavior control is central to deploying large language models (LLMs) on the web. Activation steering offers a tuning-free route to align attributes (e.g., truthfulness) that ensure trustworthy generation. Prevailing approaches…

人工智能 · 计算机科学 2025-11-19 Manjiang Yu , Hongji Li , Priyanka Singh , Xue Li , Di Wang , Lijie Hu

We introduce Mechanistic Error Reduction with Abstention (MERA), a principled framework for steering language models (LMs) to mitigate errors through selective, adaptive interventions. Unlike existing methods that rely on fixed, manually…

机器学习 · 计算机科学 2025-10-16 Anna Hedström , Salim I. Amoukou , Tom Bewley , Saumitra Mishra , Manuela Veloso

Steering intermediate representations has emerged as a powerful strategy for controlling generative models, particularly in post-deployment alignment and safety settings. However, despite its empirical success, it currently lacks a…

Controlling the behaviors of large language models (LLM) is fundamental to their safety alignment and reliable deployment. However, existing steering methods are primarily driven by empirical insights and lack theoretical performance…

机器学习 · 计算机科学 2026-05-19 Dung V. Nguyen , Hieu M. Vu , Nhi Y. Pham , Lei Zhang , Tan M. Nguyen

Network operators are generally aware of common attack vectors that they defend against. For most networks the vast majority of traffic is legitimate. However new attack vectors are continually designed and attempted by bad actors which…

机器学习 · 计算机科学 2019-04-03 Amir Ziai

Haptic guidance in a shared steering assistance system has drawn significant attention in intelligent vehicle fields, owing to its mutual communication ability for vehicle control. By exerting continuous torque on the steering wheel, both…

机器人学 · 计算机科学 2020-11-17 Zhanhong Yan , Kaiming Yang , Zheng Wang , Bo Yang , Tsutomu Kaizuka , Kimihiko Nakano

In this paper, we study the problem of `test-driving' a detector, i.e. allowing a human user to get a quick sense of how well the detector generalizes to their specific requirement. To this end, we present the first system that estimates…

计算机视觉与模式识别 · 计算机科学 2014-06-24 Rushil Anirudh , Pavan Turaga

Despite the continual advances in Advanced Driver Assistance Systems (ADAS) and the development of high-level autonomous vehicles (AV), there is a general consensus that for the short to medium term, there is a requirement for a human…

机器人学 · 计算机科学 2023-10-19 Santiago Gerling Konrad , Julie Stephany Berrio , Mao Shan , Favio Masson , Stewart Worrall

Language models can be steered by modifying their internal representations to control concepts such as emotion, style, or truthfulness in generation. However, the conditions for an effective intervention remain unclear and are often…

机器学习 · 计算机科学 2025-08-05 Jianshu She , Xinyue Li , Eric Xing , Zhengzhong Liu , Qirong Ho

As the capabilities of Vision Language Models (VLMs) continue to improve, they are increasingly targeted by jailbreak attacks. Existing defense methods face two major limitations: (1) they struggle to ensure safety without compromising the…

密码学与安全 · 计算机科学 2025-09-29 Xiyu Zeng , Siyuan Liang , Liming Lu , Haotian Zhu , Enguang Liu , Jisheng Dang , Yongbin Zhou , Shuchao Pang

Active inference is a formal approach to study cognition based on the notion that adaptive agents can be seen as engaging in a process of approximate Bayesian inference, via the minimisation of variational and expected free energies.…

人工智能 · 计算机科学 2025-08-19 Filippo Torresan , Keisuke Suzuki , Ryota Kanai , Manuel Baltieri

Language models encode task-relevant knowledge in internal representations that far exceeds their output performance, but whether mechanistic interpretability methods can bridge this knowledge-action gap has not been systematically tested.…

Policy steering is an emerging way to adapt robot behaviors at deployment-time: a learned verifier analyzes low-level action samples proposed by a pre-trained policy (e.g., diffusion policy) and selects only those aligned with the task.…

机器人学 · 计算机科学 2026-05-14 Jessie Yuan , Yilin Wu , Andrea Bajcsy

The question-answering (QA) capabilities of foundation models are highly sensitive to prompt variations, rendering their performance susceptible to superficial, non-meaning-altering changes. This vulnerability often stems from the model's…

机器学习 · 计算机科学 2024-06-07 Dyah Adila , Shuai Zhang , Boran Han , Yuyang Wang

Object detection is a critical component of a self-driving system, tasked with inferring the current states of the surrounding traffic actors. While there exist a number of studies on the problem of inferring the position and shape of…

计算机视觉与模式识别 · 计算机科学 2020-11-09 Henggang Cui , Fang-Chieh Chou , Jake Charland , Carlos Vallespi-Gonzalez , Nemanja Djuric

Is it possible to understand the intricacies of a dynamical system not solely from its input/output pattern, but also by observing the behavior of other systems within the same class? This central question drives the study presented in this…

系统与控制 · 电气工程与系统科学 2023-12-21 Marco Forgione , Filippo Pura , Dario Piga