中文
相关论文

相关论文: BeaverTails: Towards Improved Safety Alignment of …

200 篇论文

Large language models (LLMs) have a transformative impact on a variety of scientific tasks across disciplines including biology, chemistry, medicine, and physics. However, ensuring the safety alignment of these models in scientific research…

This paper presents a systematic evaluation of Large Language Models' (LLMs) behavior on long-tail distributed (encrypted) texts and their safety implications. We introduce a two-dimensional framework for assessing LLM safety: (1)…

计算与语言 · 计算机科学 2025-06-05 Utsav Maskey , Mark Dras , Usman Naseem

Large Language Models (LLMs) have been widely explored in educational scenarios. We identify a critical vulnerability in current educational LLMs, pedagogical jailbreaks, where students use answer-inducing prompts to elicit solutions rather…

计算与语言 · 计算机科学 2026-04-30 Sihang Zhao , Kangrui Yu , Youliang Yuan , Pinjia He , Hongyi Wen

We introduce a multi-turn benchmark for evaluating personalised alignment in LLM-based AI assistants, focusing on their ability to handle user-provided safety-critical contexts. Our assessment of ten leading models across five scenarios…

人机交互 · 计算机科学 2025-01-31 Lize Alberts , Benjamin Ellis , Andrei Lupu , Jakob Foerster

Alignment tuning has enabled large language models to excel in reasoning, instruction-following, and minimizing harmful generations. However, despite their widespread deployment, these models exhibit a monolingual bias, raising concerns…

计算与语言 · 计算机科学 2025-04-04 Nikhil Verma , Manasa Bharadwaj

Current open-domain conversational models can easily be made to talk in inadequate ways. Online learning from conversational feedback given by the conversation partner is a promising avenue for a model to improve and adapt, so as to…

计算与语言 · 计算机科学 2022-05-06 Megan Ung , Jing Xu , Y-Lan Boureau

Optimizing large language models (LLMs) for downstream use cases often involves the customization of pre-trained LLMs through further fine-tuning. Meta's open release of Llama models and OpenAI's APIs for fine-tuning GPT-3.5 Turbo on custom…

计算与语言 · 计算机科学 2023-10-06 Xiangyu Qi , Yi Zeng , Tinghao Xie , Pin-Yu Chen , Ruoxi Jia , Prateek Mittal , Peter Henderson

Large language models (LLMs) are increasingly deployed as conversational assistants in open-domain, multi-turn settings, where users often provide incomplete or ambiguous information. However, existing LLM-focused clarification benchmarks…

计算与语言 · 计算机科学 2025-12-25 Sichun Luo , Yi Huang , Mukai Li , Shichang Meng , Fengyuan Liu , Zefa Hu , Junlan Feng , Qi Liu

Natural Question Answering (QA) datasets play a crucial role in evaluating the capabilities of large language models (LLMs), ensuring their effectiveness in real-world applications. Despite the numerous QA datasets that have been developed…

Large Language Models (LLMs) have the potential to significantly enhance threat intelligence by automating the collection, preprocessing, and analysis of threat data. However, the usability of these tools is critical to ensure their…

密码学与安全 · 计算机科学 2024-09-24 Sanchana Srikanth , Mohammad Hasanuzzaman , Farah Tasnur Meem

Multimodal Large Language Models (MLLMs) pose critical safety challenges, as they are susceptible not only to adversarial attacks such as jailbreaking but also to inadvertently generating harmful content for benign users. While internal…

机器学习 · 计算机科学 2026-03-17 Ming Wen , Kun Yang , Xin Chen , Jingyu Zhang , Dingding Han , Shiwen Cui , Yuedong Xu

The safety of Large Language Models (LLMs) has gained increasing attention in recent years, but there still lacks a comprehensive approach for detecting safety issues within LLMs' responses in an aligned, customizable and explainable…

计算与语言 · 计算机科学 2024-11-06 Zhexin Zhang , Yida Lu , Jingyuan Ma , Di Zhang , Rui Li , Pei Ke , Hao Sun , Lei Sha , Zhifang Sui , Hongning Wang , Minlie Huang

Studying the robustness of Large Language Models (LLMs) to unsafe behaviors is an important topic of research today. Building safety classification models or guard models, which are fine-tuned models for input/output safety classification…

计算与语言 · 计算机科学 2025-07-30 Sowmya Vajjala

Large language models (LLMs) have shown great potential as general-purpose AI assistants in various domains. To meet the requirements of different applications, LLMs are often customized by further fine-tuning. However, the powerful…

机器学习 · 计算机科学 2023-11-07 Xin Zhou , Yi Lu , Ruotian Ma , Tao Gui , Qi Zhang , Xuanjing Huang

The application scope of Large Language Models (LLMs) continues to expand, leading to increasing interest in personalized LLMs that align with human values. However, aligning these models with individual values raises significant safety…

计算与语言 · 计算机科学 2025-06-10 Sooyung Choi , Jaehyeok Lee , Xiaoyuan Yi , Jing Yao , Xing Xie , JinYeong Bak

As Large Language Models (LLMs) continue to advance in understanding and generating long sequences, new safety concerns have been introduced through the long context. However, the safety of LLMs in long-context tasks remains under-explored,…

计算与语言 · 计算机科学 2025-02-25 Yida Lu , Jiale Cheng , Zhexin Zhang , Shiyao Cui , Cunxiang Wang , Xiaotao Gu , Yuxiao Dong , Jie Tang , Hongning Wang , Minlie Huang

Open-weight large language models (LLMs) unlock huge benefits in innovation, personalization, privacy, and democratization. However, their core advantage - modifiability - opens the door to systemic risks: bad actors can trivially subvert…

计算机与社会 · 计算机科学 2025-07-17 Ann-Kathrin Dombrowski , Dillon Bowen , Adam Gleave , Chris Cundy

Predicting user behavior is essential for intelligent assistant services, yet deep learning models often struggle to capture long-tailed behaviors. Large language models (LLMs), with their pretraining on vast corpora containing rich…

计算与语言 · 计算机科学 2026-04-14 Fanjin Meng , Jingtao Ding , Jiahui Gong , Chen Yang , Hong Chen , Zuojian Wang , Haisheng Lu , Yong Li

Large language models (LLMs) have been widely applied in various fields due to their excellent capability for memorizing knowledge and chain of thought (CoT). When these language models are applied in the field of psychological counseling,…

计算与语言 · 计算机科学 2023-11-02 Yirong Chen , Xiaofen Xing , Jingkai Lin , Huimin Zheng , Zhenyu Wang , Qi Liu , Xiangmin Xu

Mental health is a growing global concern, prompting interest in AI-driven solutions to expand access to psychosocial support. \emph{Peer support}, grounded in lived experience, offers a valuable complement to professional care. However,…

人机交互 · 计算机科学 2026-04-10 Kellie Yu Hui Sim , Roy Ka-Wei Lee , Kenny Tsu Wei Choo