中文
相关论文

相关论文: Beyond the Safety Bundle: Auditing the Helpful and…

200 篇论文

Large Language Models (LLMs) interact with millions of people worldwide in applications such as customer support, education and healthcare. However, their ability to produce deceptive outputs, whether intentionally or inadvertently, poses…

计算与语言 · 计算机科学 2025-10-17 Marwa Abdulhai , Ryan Cheng , Aryansh Shrivastava , Natasha Jaques , Yarin Gal , Sergey Levine

As large language models (LLMs) have increased in their capabilities, so does their potential for dual use. To reduce harmful outputs, produces and vendors of LLMs have used reinforcement learning with human feedback (RLHF). In tandem, LLM…

计算与语言 · 计算机科学 2024-04-09 Qiusi Zhan , Richard Fang , Rohan Bindu , Akul Gupta , Tatsunori Hashimoto , Daniel Kang

Nowadays, developers increasingly rely on solutions powered by Large Language Models (LLM) to assist them with their coding tasks. This makes it crucial to align these tools with human values to prevent malicious misuse. In this paper, we…

软件工程 · 计算机科学 2025-04-03 Ali Al-Kaswan , Sebastian Deatc , Begüm Koç , Arie van Deursen , Maliheh Izadi

Human feedback can alter language models in unpredictable and undesirable ways, as practitioners lack a clear understanding of what feedback data encodes. While prior work studies preferences over certain attributes (e.g., length or…

计算与语言 · 计算机科学 2026-04-14 Rajiv Movva , Smitha Milli , Sewon Min , Emma Pierson

With the widespread availability of LLMs since the release of ChatGPT and increased public scrutiny, commercial model development appears to have focused their efforts on 'safety' training concerning legal liabilities at the expense of…

计算与语言 · 计算机科学 2024-08-02 Alina Leidinger , Richard Rogers

Large language models (LLMs) excel in diverse applications but face dual challenges: generating harmful content under jailbreak attacks and over-refusal of benign queries due to rigid safety mechanisms. These issues are further complicated…

人工智能 · 计算机科学 2025-11-04 Yifan Xia , Guorui Chen , Wenqian Yu , Zhijiang Li , Philip Torr , Jindong Gu

Large Language Models (LLM) have made remarkable progress, but concerns about potential biases and harmful content persist. To address these apprehensions, we introduce a practical solution for ensuring LLM's safe and ethical use. Our novel…

密码学与安全 · 计算机科学 2025-04-24 Chaima Njeh , Haïfa Nakouri , Fehmi Jaafar

Large language models (LLMs) increasingly operate on long inputs, yet their behavior when harmful sentences are sparsely embedded within such inputs remains poorly understood. We present a sensitivity analysis that probes how LLMs extract…

计算与语言 · 计算机科学 2026-05-27 Faeze Ghorbanpour , Alexander Fraser

Large language models are being deployed as mental health support agents at scale, yet only 16% of LLM-based chatbot interventions have undergone rigorous clinical efficacy testing, and simulations reveal psychological deterioration in over…

计算与语言 · 计算机科学 2026-04-28 Suhas BN , Andrew M. Sherrill , Rosa I. Arriaga , Chris W. Wiese , Saeed Abdullah

Large language models (LLMs) are increasingly deployed in contexts where their failures can have direct sociopolitical consequences. Yet, existing safety benchmarks rarely test vulnerabilities in domains such as political manipulation,…

计算与语言 · 计算机科学 2026-02-24 Punya Syon Pandey , Hai Son Le , Devansh Bhardwaj , Rada Mihalcea , Zhijing Jin

The growing demand for accessible mental health support, compounded by workforce shortages and logistical barriers, has led to increased interest in utilizing Large Language Models (LLMs) for scalable and real-time assistance. However,…

Reinforcement Learning from Human Feedback (RLHF) has been credited as the key advance that has allowed Large Language Models (LLMs) to effectively follow instructions and produce useful assistance. Classically, this involves generating…

机器学习 · 计算机科学 2024-02-02 Alex J. Chan , Hao Sun , Samuel Holt , Mihaela van der Schaar

This paper presents a systematic evaluation of Large Language Models' (LLMs) behavior on long-tail distributed (encrypted) texts and their safety implications. We introduce a two-dimensional framework for assessing LLM safety: (1)…

计算与语言 · 计算机科学 2025-06-05 Utsav Maskey , Mark Dras , Usman Naseem

Large language models (LLMs) are increasingly applied in mental health support systems, where reliable recognition of high-risk states such as suicidal ideation and self-harm is safety-critical. However, existing evaluations primarily rely…

人工智能 · 计算机科学 2026-03-12 Yihe Zhang , Cheyenne N Mohawk , Kaiying Han , Vijay Srinivas Tida , Manyu Li , Xiali Hei

Large language models (LLMs) are increasingly deployed as tool-using agents, shifting safety concerns from harmful text generation to harmful task completion. Deployed systems often condition on user profiles or persistent memory, yet agent…

人工智能 · 计算机科学 2026-03-18 Caglar Yildirim

We study the problem of reinforcement learning from human feedback (RLHF), a critical problem in training large language models, from a theoretical perspective. Our main contribution is the design of novel sample-efficient RLHF algorithms…

机器学习 · 计算机科学 2025-08-11 Han Qi , Haochen Yang , Qiaosheng Zhang , Zhuoran Yang

Time series neural networks perform exceptionally well in real-world applications but encounter challenges such as limited scalability, poor generalization, and suboptimal zero-shot performance. Inspired by large language models, there is…

机器学习 · 计算机科学 2025-01-28 Yongzhi Qi , Hao Hu , Dazhou Lei , Jianshen Zhang , Zhengxin Shi , Yulin Huang , Zhengyu Chen , Xiaoming Lin , Zuo-Jun Max Shen

While Reinforcement Learning from Human Feedback (RLHF) is widely used to align Large Language Models (LLMs) with human preferences, it typically assumes homogeneous preferences across users, overlooking diverse human values and minority…

计算与语言 · 计算机科学 2025-10-28 Yijiang River Dong , Tiancheng Hu , Yinhong Liu , Ahmet Üstün , Nigel Collier

Large language models are increasingly used for creative writing and engagement content, raising safety concerns about the outputs. Therefore, casting humor generation as a testbed, this work evaluates how funniness optimization in modern…

计算与语言 · 计算机科学 2025-10-22 Atharvan Dogra , Soumya Suvra Ghosal , Ameet Deshpande , Ashwin Kalyan , Dinesh Manocha

Reinforcement learning from human feedback (RLHF) is a powerful technique for training agents to perform difficult-to-specify tasks. However, human feedback can be noisy, particularly when human teachers lack relevant knowledge or…

机器学习 · 计算机科学 2022-11-15 Oliver Daniels-Koch , Rachel Freedman
‹ 上一页 1 8 9 10 下一页 ›