中文
相关论文

相关论文: Reverse Constitutional AI: A Framework for Control…

200 篇论文

Traditional methods for aligning Large Language Models (LLMs), such as Reinforcement Learning from Human Feedback (RLHF) and Direct Preference Optimization (DPO), rely on implicit principles, limiting interpretability. Constitutional AI…

机器学习 · 计算机科学 2025-04-01 Carl-Leander Henneking , Claas Beger

Large Language Models have emerged as transformative tools for Security Operations Centers, enabling automated log analysis, phishing triage, and malware explanation; however, deployment in adversarial cybersecurity environments exposes…

密码学与安全 · 计算机科学 2026-01-13 Mohammed Himayath Ali , Mohammed Aqib Abdullah , Mohammed Mudassir Uddin , Shahnawaz Alam

Feedback data is widely used for fine-tuning and evaluating state-of-the-art AI models. Pairwise text preferences, where human or AI annotators select the "better" of two options, are particularly common. Such preferences are used to train…

计算与语言 · 计算机科学 2025-04-22 Arduin Findeis , Timo Kaufmann , Eyke Hüllermeier , Samuel Albanie , Robert Mullins

Constitutional AI is a method to oversee and control LLMs based on a set of rules written in natural language. These rules are typically written by human experts, but could in principle be learned automatically given sufficient training…

人工智能 · 计算机科学 2026-03-18 Rushil Thareja , Gautam Gupta , Francesco Pinto , Nils Lukas

With the rapid development of large language models (LLMs), aligning LLMs with human values and societal norms to ensure their reliability and safety has become crucial. Reinforcement learning with human feedback (RLHF) and Constitutional…

计算与语言 · 计算机科学 2024-03-28 Xiusi Chen , Hongzhi Wen , Sreyashi Nag , Chen Luo , Qingyu Yin , Ruirui Li , Zheng Li , Wei Wang

As language models continue to grow larger, the cost of acquiring high-quality training data has increased significantly. Collecting human feedback is both expensive and time-consuming, and manual labels can be noisy, leading to an…

人工智能 · 计算机科学 2025-04-08 Xue Zhang

A crucial consideration when developing and deploying Large Language Models (LLMs) is the human values to which these models are aligned. In the constitutional framework of alignment models are aligned to a set of principles (the…

机器学习 · 计算机科学 2026-01-27 Henry Bell , Lara Neubauer da Costa Schertel , Bochu Ding , Brandon Fain

There is growing consensus that language model (LM) developers should not be the sole deciders of LM behavior, creating a need for methods that enable the broader public to collectively shape the behavior of LM systems that affect them. To…

人工智能 · 计算机科学 2024-06-13 Saffron Huang , Divya Siddarth , Liane Lovitt , Thomas I. Liao , Esin Durmus , Alex Tamkin , Deep Ganguli

Safely aligning large language models (LLMs) often demands extensive human-labeled preference data, a process that's both costly and time-consuming. While synthetic data offers a promising alternative, current methods frequently rely on…

密码学与安全 · 计算机科学 2025-06-13 Kyubyung Chae , Hyunbin Jin , Taesup Kim

Recent research has increasingly focused on training large language models (LLMs) using federated learning, known as FedLLM. However, responsible AI (RAI), which aims to ensure safe and trustworthy responses, remains underexplored in this…

计算与语言 · 计算机科学 2026-05-19 Eunchung Noh , Jeonghun Baek

Constitutional AI (CAI) guides LLM behavior using constitutions, but identifying which principles are most effective for model alignment remains an open challenge. We introduce the C3AI framework (\textit{Crafting Constitutions for CAI…

人工智能 · 计算机科学 2025-02-25 Yara Kyrychenko , Ke Zhou , Edyta Bogucka , Daniele Quercia

Recent incidents highlight safety risks in Large Language Models (LLMs), motivating research into alignment methods like Constitutional AI (CAI). This paper explores CAI's self-critique mechanism on small, uncensored 7-9B parameter models:…

机器学习 · 计算机科学 2025-04-14 Antonio-Gabriel Chacón Menke , Phan Xuan Tan

Retrieval-Augmented Generation (RAG) enhances large language models (LLMs) by incorporating external knowledge, but its openness introduces vulnerabilities that can be exploited by poisoning attacks. Existing poisoning methods for RAG…

密码学与安全 · 计算机科学 2025-05-27 Chunyang Li , Junwei Zhang , Anda Cheng , Zhuo Ma , Xinghua Li , Jianfeng Ma

Generative AI, in particular text-based "foundation models" (large models trained on a huge variety of information including the internet), can generate speech that could be problematic under a wide range of liability regimes. Machine…

计算机与社会 · 计算机科学 2023-08-21 Peter Henderson , Tatsunori Hashimoto , Mark Lemley

AI-enabled Security Orchestration, Automation, and Response (SOAR) systems increasingly employ autonomous agents for cyber defense, yet their resilience to adaptive adversaries is underexplored. We introduce an autonomous red teaming…

密码学与安全 · 计算机科学 2026-05-19 Ayan Javeed Shaikh , Nathaniel D. Bastian , Ankit Shah

Large Language Models (LLMs) are being integrated into professional domains, yet their limitations in such high-stakes fields as law remain poorly understood. In response, this paper introduces examples of critical challenges to the…

人工智能 · 计算机科学 2026-01-27 Eljas Linna , Tuula Linna

The dual offensive and defensive utility of Large Language Models (LLMs) highlights a critical gap in AI security: the lack of unified frameworks for dynamic, iterative adversarial adaptation hardening. To bridge this gap, we propose the…

密码学与安全 · 计算机科学 2026-01-28 Lige Huang , Zicheng Liu , Jie Zhang , Lewen Yan , Dongrui Liu , Jing Shao

As frontier AI models are deployed in high-stakes decision pipelines, their ability to maintain metacognitive stability (knowing what they do not know, detecting errors, seeking clarification) under adversarial pressure is a critical safety…

人工智能 · 计算机科学 2026-05-15 Rahul Kumar
‹ 上一页 1 2 3 10 下一页 ›