中文
相关论文

相关论文: Value Alignment Tax: Measuring Value Trade-offs in…

200 篇论文

Behavioral evaluation is the dominant paradigm for assessing alignment in large language models (LLMs). In current practice, observed compliance under finite evaluation protocols is treated as evidence of latent alignment. However, the…

机器学习 · 计算机科学 2026-02-10 Igor Santos-Grueiro

Aligning AI agents with human values is challenging due to diverse and subjective notions of values. Standard alignment methods often aggregate crowd feedback, which can result in the suppression of unique or minority preferences. We…

人工智能 · 计算机科学 2024-10-30 Carter Blair , Kate Larson , Edith Law

Do Large Language Models (LLMs) hold positions that conflict with your country's values? Occasionally they do! However, existing works primarily focus on ethical reviews, failing to capture the diversity of national values, which encompass…

计算与语言 · 计算机科学 2025-04-22 Weijie Shi , Chengyi Ju , Chengzhong Liu , Jiaming Ji , Jipeng Zhang , Ruiyuan Zhang , Jia Zhu , Jiajie Xu , Yaodong Yang , Sirui Han , Yike Guo

The construction of high-quality parallel corpora for translation research has increasingly evolved from simple sentence alignment to complex, multi-layered annotation tasks. This methodological shift presents significant challenges for…

计算与语言 · 计算机科学 2026-02-12 Baorong Huang , Ali Asiri

Detecting anomalies in time series data is crucial for finance, healthcare, sensor networks, and industrial monitoring applications. However, time series anomaly detection often suffers from sparse labels, complex temporal patterns, and…

机器学习 · 计算机科学 2026-01-07 Bahareh Golchin , Banafsheh Rekabdar , Danielle Justo

Despite advances in large language models (LLMs) on reasoning and instruction-following tasks, it is unclear whether they can reliably produce outputs aligned with a variety of user goals, a concept called steerability. Two gaps in current…

计算与语言 · 计算机科学 2026-01-21 Trenton Chang , Tobias Schnabel , Adith Swaminathan , Jenna Wiens

The rise of LLM-based agents has opened new frontiers in AI applications, yet evaluating these agents remains a complex and underdeveloped area. This survey provides an in-depth overview of the emerging field of LLM agent evaluation,…

机器学习 · 计算机科学 2025-07-30 Mahmoud Mohammadi , Yipeng Li , Jane Lo , Wendy Yip

Visual content and accompanied audio signals naturally formulate a joint representation to improve audio-visual (AV) related applications. While studies develop various AV representation learning frameworks, the importance of AV data…

计算机视觉与模式识别 · 计算机科学 2024-11-01 Shentong Mo , Yibing Song

The rapid progress in Large Language Models (LLMs) poses potential risks such as generating unethical content. Assessing LLMs' values can help expose their misalignment, but relies on reference-free evaluators, e.g., fine-tuned LLMs or…

计算与语言 · 计算机科学 2024-07-16 Jing Yao , Xiaoyuan Yi , Xing Xie

Safety alignment is crucial to ensure that large language models (LLMs) behave in ways that align with human preferences and prevent harmful actions during inference. However, recent studies show that the alignment can be easily compromised…

机器学习 · 计算机科学 2024-11-01 ShengYun Peng , Pin-Yu Chen , Matthew Hull , Duen Horng Chau

Improving the alignment of Large Language Models (LLMs) with respect to the cultural values that they encode has become an increasingly important topic. In this work, we study whether we can exploit existing knowledge about cultural values…

计算与语言 · 计算机科学 2025-09-09 Rochelle Choenni , Ekaterina Shutova

Pair trading is one of the most discussed topics among financial researches. Despite a growing base of work, portfolio management for multivariate time series is rarely discussed. On the other hand, most researches focus on refining…

投资组合管理 · 定量金融 2021-06-18 Chenyanzi Yu , Tianyang Xie

A default assumption in reinforcement learning (RL) and optimal control is that observations arrive at discrete time points on a fixed clock cycle. Yet, many applications involve continuous-time systems where the time discretization, in…

Research on the 'cultural alignment' of Large Language Models (LLMs) has emerged in response to growing interest in understanding representation across diverse stakeholders. Current approaches to evaluating cultural alignment through…

计算机与社会 · 计算机科学 2025-04-10 Ariba Khan , Stephen Casper , Dylan Hadfield-Menell

Language embeds information about social, cultural, and political values people hold. Prior work has explored social and potentially harmful biases encoded in Pre-Trained Language models (PTLMs). However, there has been no systematic study…

计算与语言 · 计算机科学 2025-08-29 Arnav Arora , Lucie-Aimée Kaffee , Isabelle Augenstein

This work shows that value-aware model learning, known for its numerous theoretical benefits, is also practically viable for solving challenging continuous control tasks in prevalent model-based reinforcement learning algorithms. First, we…

机器学习 · 计算机科学 2022-01-31 Nirbhay Modhe , Harish Kamath , Dhruv Batra , Ashwin Kalyan

A common assumption in machine learning is that samples are independently and identically distributed (i.i.d). However, the contributions of different samples are not identical in training. Some samples are difficult to learn and some…

机器学习 · 计算机科学 2021-11-23 Ou Wu , Weiyao Zhu , Yingjun Deng , Haixiang Zhang , Qinghu Hou

Pre-trained vision-language models (VLMs) have shown impressive results in various visual classification tasks. However, we often fail to fully unleash their potential when adapting them for new concept understanding due to limited…

计算机视觉与模式识别 · 计算机科学 2024-10-08 Yuhan Zhu , Yuyang Ji , Zhiyu Zhao , Gangshan Wu , Limin Wang

Aligning AI systems with organizational decision-making is typically framed as a single-target problem: make the model behave like the organization. We argue this framing obscures a deeper pluralistic challenge. We rely on a decision-policy…

人工智能 · 计算机科学 2026-05-26 Niklas Weller , Emilio Barkett

Vision transformers (ViTs) can be trained using various learning paradigms, from fully supervised to self-supervised. Diverse training protocols often result in significantly different feature spaces, which are usually compared through…

计算机视觉与模式识别 · 计算机科学 2024-12-10 Johanna Vielhaben , Dilyara Bareeva , Jim Berend , Wojciech Samek , Nils Strodthoff