中文
相关论文

相关论文: The Ethics of LLM Sandbox and Persona Dynamics

200 篇论文

Large language models (LLMs) have become increasingly sophisticated, leading to widespread deployment in sensitive applications where safety and reliability are paramount. However, LLMs have inherent risks accompanying them, including bias,…

密码学与安全 · 计算机科学 2024-06-21 Suriya Ganesh Ayyamperumal , Limin Ge

The AI era has ushered in Large Language Models (LLM) to the technological forefront, which has been much of the talk in 2023, and is likely to remain as such for many years to come. LLMs are the AI models that are the power house behind…

密码学与安全 · 计算机科学 2026-01-22 Anjanava Biswas , Wrick Talukdar

Large language models (LLMs) rely on safety alignment to avoid responding to malicious user inputs. Unfortunately, jailbreak can circumvent safety guardrails, resulting in LLMs generating harmful content and raising concerns about LLM…

计算与语言 · 计算机科学 2024-06-14 Zhenhong Zhou , Haiyang Yu , Xinghua Zhang , Rongwu Xu , Fei Huang , Yongbin Li

Recent advances in large language models (LLMs) have enabled their use in complex agentic roles, involving decision-making with humans or other agents, making ethical alignment a key AI safety concern. While prior work has examined both…

Recent advances in large language models (LLMs) have led to increasingly sophisticated safety protocols and features designed to prevent harmful, unethical, or unauthorized outputs. However, these guardrails remain susceptible to novel and…

计算与语言 · 计算机科学 2025-07-08 Annika M Schoene , Cansu Canca

Large Language Models (LLMs) exhibit surprisingly diverse risk preferences when acting as AI decision makers, a crucial characteristic whose origins remain poorly understood despite their expanding economic roles. We analyze 50 LLMs using…

综合经济学 · 经济学 2025-06-11 Shumiao Ouyang , Hayong Yun , Xingjian Zheng

In this study, we tackle a growing concern around the safety and ethical use of large language models (LLMs). Despite their potential, these models can be tricked into producing harmful or unethical content through various sophisticated…

计算与语言 · 计算机科学 2024-11-19 Somnath Banerjee , Sayan Layek , Rima Hazra , Animesh Mukherjee

The rapid advancement of large language models (LLMs) raises critical concerns about their ethical alignment, particularly in scenarios where human and AI co-exist under the conflict of interest. This work introduces an extendable,…

人机交互 · 计算机科学 2025-05-29 Zhihong Chen , Yiqian Yang , Jinzhao Zhou , Qiang Zhang , Chin-Teng Lin , Yiqun Duan

In the burgeoning field of Large Language Models (LLMs), developing a robust safety mechanism, colloquially known as "safeguards" or "guardrails", has become imperative to ensure the ethical use of LLMs within prescribed boundaries. This…

密码学与安全 · 计算机科学 2024-06-06 Yi Dong , Ronghui Mu , Yanghao Zhang , Siqi Sun , Tianle Zhang , Changshun Wu , Gaojie Jin , Yi Qi , Jinwei Hu , Jie Meng , Saddek Bensalem , Xiaowei Huang

As artificial intelligence systems become increasingly agentic, capable of general reasoning, planning, and value prioritization, current safety practices that treat obedience as a proxy for ethical behavior are becoming inadequate. This…

人工智能 · 计算机科学 2025-07-04 Joseph Boland

Current LLMs are trained to refuse potentially harmful input queries regardless of whether users actually had harmful intents, causing a tradeoff between safety and user experience. Through a study of 480 participants evaluating 3,840…

计算与语言 · 计算机科学 2025-12-04 Mingqian Zheng , Wenjia Hu , Patrick Zhao , Motahhare Eslami , Jena D. Hwang , Faeze Brahman , Carolyn Rose , Maarten Sap

Recent reports of large language models (LLMs) exhibiting behaviors such as deception, threats, or blackmail are often interpreted as evidence of alignment failure or emergent malign agency. We argue that this interpretation rests on a…

人工智能 · 计算机科学 2026-01-14 Didier Sornette , Sandro Claudio Lera , Ke Wu

LLM-enabled robots prioritizing scarce assistance in social settings face pluralistic values and LLM behavioral variability: reasonable people can disagree about who is helped first, while LLM-mediated interaction policies vary across…

人工智能 · 计算机科学 2026-03-18 Carmen Ng

This paper explores the development of an ethical guardrail framework for AI systems, emphasizing the importance of customizable guardrails that align with diverse user values and underlying ethics. We address the challenges of AI ethics by…

计算机与社会 · 计算机科学 2024-11-25 Kristina Šekrst , Jeremy McHugh , Jonathan Rodriguez Cefalu

Advancements in large language models (LLMs) have renewed concerns about AI alignment - the consistency between human and AI goals and values. As various jurisdictions enact legislation on AI safety, the concept of alignment must be defined…

计算机与社会 · 计算机科学 2025-02-26 Claudia Biancotti , Carolina Camassa , Andrea Coletta , Oliver Giudice , Aldo Glielmo

The recent rise in popularity of large language models (LLMs) has prompted considerable concerns about their moral capabilities. Although considerable effort has been dedicated to aligning LLMs with human moral values, existing benchmarks…

人工智能 · 计算机科学 2025-08-19 Alessio Galatolo , Luca Alberto Rappuoli , Katie Winkle , Meriem Beloucif

Large language models (LLMs) are useful tools with the capacity for performing specific types of knowledge work at an effective scale. However, LLM deployments in high-risk and safety-critical domains pose unique challenges, notably the…

This paper comprehensively explores the ethical challenges arising from security threats to Large Language Models (LLMs). These intricate digital repositories are increasingly integrated into our daily lives, making them prime targets for…

密码学与安全 · 计算机科学 2024-07-11 Ashutosh Kumar , Shiv Vignesh Murthy , Sagarika Singh , Swathy Ragupathy

The rapid progress in Large Language Models (LLMs) could transform many fields, but their fast development creates significant challenges for oversight, ethical creation, and building user trust. This comprehensive review looks at key trust…

计算机与社会 · 计算机科学 2024-07-22 Md Meftahul Ferdaus , Mahdi Abdelguerfi , Elias Ioup , Kendall N. Niles , Ken Pathak , Steven Sloan

This study addresses ethical issues surrounding Large Language Models (LLMs) within the field of artificial intelligence. It explores the common ethical challenges posed by both LLMs and other AI systems, such as privacy and fairness, as…

计算机与社会 · 计算机科学 2025-06-17 Junfeng Jiao , Saleh Afroogh , Yiming Xu , Connor Phillips
‹ 上一页 1 2 3 10 下一页 ›