中文
相关论文

相关论文: AM$^3$Safety: Towards Data Efficient Alignment of …

200 篇论文

Humans are prone to cognitive distortions -- biased thinking patterns that lead to exaggerated responses to specific stimuli, albeit in very different contexts. This paper demonstrates that advanced Multimodal Large Language Models (MLLMs)…

计算与语言 · 计算机科学 2024-06-27 Xirui Li , Hengguang Zhou , Ruochen Wang , Tianyi Zhou , Minhao Cheng , Cho-Jui Hsieh

We introduce SafeWork-R1, a cutting-edge multimodal reasoning model that demonstrates the coevolution of capabilities and safety. It is developed by our proposed SafeLadder framework, which incorporates large-scale, progressive,…

人工智能 · 计算机科学 2025-08-08 Shanghai AI Lab , : , Yicheng Bao , Guanxu Chen , Mingkang Chen , Yunhao Chen , Chiyu Chen , Lingjie Chen , Sirui Chen , Xinquan Chen , Jie Cheng , Yu Cheng , Dengke Deng , Yizhuo Ding , Dan Ding , Xiaoshan Ding , Yi Ding , Zhichen Dong , Lingxiao Du , Yuyu Fan , Xinshun Feng , Yanwei Fu , Yuxuan Gao , Ruijun Ge , Tianle Gu , Lujun Gui , Jiaxuan Guo , Qianxi He , Yuenan Hou , Xuhao Hu , Hong Huang , Kaichen Huang , Shiyang Huang , Yuxian Jiang , Shanzhe Lei , Jie Li , Lijun Li , Hao Li , Juncheng Li , Xiangtian Li , Yafu Li , Lingyu Li , Xueyan Li , Haotian Liang , Dongrui Liu , Qihua Liu , Zhixuan Liu , Bangwei Liu , Huacan Liu , Yuexiao Liu , Zongkai Liu , Chaochao Lu , Yudong Lu , Xiaoya Lu , Zhenghao Lu , Qitan Lv , Caoyuan Ma , Jiachen Ma , Xiaoya Ma , Zhongtian Ma , Lingyu Meng , Ziqi Miao , Yazhe Niu , Yuezhang Peng , Yuan Pu , Han Qi , Chen Qian , Xingge Qiao , Jingjing Qu , Jiashu Qu , Wanying Qu , Wenwen Qu , Xiaoye Qu , Qihan Ren , Qingnan Ren , Qingyu Ren , Jing Shao , Wenqi Shao , Shuai Shao , Dongxing Shi , Xin Song , Xinhao Song , Yan Teng , Xuan Tong , Yingchun Wang , Xuhong Wang , Shujie Wang , Xin Wang , Yige Wang , Yixu Wang , Yuanfu Wang , Futing Wang , Ruofan Wang , Wenjie Wang , Yajie Wang , Muhao Wei , Xiaoyu Wen , Fenghua Weng , Yuqi Wu , Yingtong Xiong , Xingcheng Xu , Chao Yang , Yue Yang , Yang Yao , Yulei Ye , Zhenyun Yin , Yi Yu , Bo Zhang , Qiaosheng Zhang , Jinxuan Zhang , Yexin Zhang , Yinqiang Zheng , Hefeng Zhou , Zhanhui Zhou , Pengyu Zhu , Qingzi Zhu , Yubo Zhu , Bowen Zhou

Recent studies reveal that vision-language models (VLMs) become more susceptible to harmful requests and jailbreak attacks after integrating the vision modality, exhibiting greater vulnerability than their text-only LLM backbones. To…

计算机视觉与模式识别 · 计算机科学 2025-02-19 Xiaohan Zou , Jian Kang , George Kesidis , Lu Lin

In recent years, safety risks associated with large language models have become increasingly prominent, highlighting the urgent need to mitigate the generation of toxic and harmful content. The mainstream paradigm for LLM safety alignment…

Large Language Models (LLMs) are increasingly used as high level controllers for autonomous Unmanned Aerial Vehicle (UAV) missions. However, existing evaluations rarely assess whether such agents remain safe, protocol compliant, and…

系统与控制 · 电气工程与系统科学 2026-01-08 Mohamed Amine Ferrag , Abderrahmane Lakas , Merouane Debbah

Recent studies on the safety alignment of large language models (LLMs) have revealed that existing approaches often operate superficially, leaving models vulnerable to various adversarial attacks. Despite their significance, these studies…

密码学与安全 · 计算机科学 2025-06-02 Jianwei Li , Jung-Eun Kim

Aligning large language models (LLMs) with human values and safety constraints is challenging, especially when objectives like helpfulness, truthfulness, and avoidance of harm conflict. Reinforcement Learning from Human Feedback (RLHF) has…

计算与语言 · 计算机科学 2025-03-31 Xuying Li , Zhuo Li , Yuji Kosuga , Victor Bian

Achieving robust safety alignment in large language models (LLMs) while preserving their utility remains a fundamental challenge. Existing approaches often struggle to balance comprehensive safety with fine-grained controllability at the…

人工智能 · 计算机科学 2025-09-25 Huizhen Shu , Xuying Li , Zhuo Li

Large language models (LLMs), despite possessing latent safety understanding from their vast pretraining data, remain vulnerable to generating harmful content and exhibit issues such as over-refusal and utility degradation after safety…

人工智能 · 计算机科学 2025-07-22 Yi Zhang , An Zhang , XiuYu Zhang , Leheng Sheng , Yuxin Chen , Zhenkai Liang , Xiang Wang

Large Language Models (LLMs) are powerful tools for modern applications, but their computational demands limit accessibility. Quantization offers efficiency gains, yet its impact on safety and trustworthiness remains poorly understood. To…

密码学与安全 · 计算机科学 2025-07-01 Artyom Kharinaev , Viktor Moskvoretskii , Egor Shvetsov , Kseniia Studenikina , Bykov Mikhail , Evgeny Burnaev

Safety lies at the core of developing and deploying large language models (LLMs). However, previous safety benchmarks only concern the safety in one language, e.g. the majority language in the pretraining data such as English. In this work,…

计算与语言 · 计算机科学 2024-06-21 Wenxuan Wang , Zhaopeng Tu , Chang Chen , Youliang Yuan , Jen-tse Huang , Wenxiang Jiao , Michael R. Lyu

While reasoning models have achieved remarkable success in complex reasoning tasks, their increasing power necessitates stringent safety measures. For safety alignment, the core challenge lies in the inherent trade-off between safety and…

密码学与安全 · 计算机科学 2026-02-17 Yanbo Wang , Minzheng Wang , Jian Liang , Lu Wang , Yongcan Yu , Ran He

Fine-tuning large language models (LLMs) on additional datasets is often necessary to optimize them for specific downstream tasks. However, existing safety alignment measures, which restrict harmful behavior during inference, are…

计算与语言 · 计算机科学 2024-10-15 Minjun Zhu , Linyi Yang , Yifan Wei , Ningyu Zhang , Yue Zhang

The rapid advancement of multi-modal large reasoning models (MLRMs) -- enhanced versions of multimodal language models (MLLMs) equipped with reasoning capabilities -- has revolutionized diverse applications. However, their safety…

机器学习 · 计算机科学 2025-04-15 Junfeng Fang , Yukai Wang , Ruipeng Wang , Zijun Yao , Kun Wang , An Zhang , Xiang Wang , Tat-Seng Chua

Recent advancements in model architectures and length extrapolation techniques have significantly extended the context length of large language models (LLMs), paving the way for their application in increasingly complex tasks. However,…

Large Language Models (LLMs) are swiftly advancing in architecture and capability, and as they integrate more deeply into complex systems, the urgency to scrutinize their security properties grows. This paper surveys research in the…

计算与语言 · 计算机科学 2023-10-18 Erfan Shayegani , Md Abdullah Al Mamun , Yu Fu , Pedram Zaree , Yue Dong , Nael Abu-Ghazaleh

Large Language Models (LLMs) have demonstrated impressive capabilities in various tasks, including instruction following, which is crucial for aligning model outputs with user expectations. However, evaluating LLMs' ability to follow…

Vision Language Models (VLMs) have become essential backbones for multimodal intelligence, yet significant safety challenges limit their real-world application. While textual inputs are often effectively safeguarded, adversarial visual…

计算机视觉与模式识别 · 计算机科学 2025-02-11 Yi Ding , Bolian Li , Ruqi Zhang

Integrated Speech and Large Language Models (SLMs) that can follow speech instructions and generate relevant text responses have gained popularity lately. However, the safety and robustness of these models remains largely unclear. In this…

As Multimodal Large Language Models (MLLMs) become an indispensable assistant in human life, the unsafe content generated by MLLMs poses a danger to human behavior, perpetually overhanging human society like a sword of Damocles. To…

计算与语言 · 计算机科学 2026-04-21 Xinyue Lou , Jinan Xu , Jingyi Yin , Xiaolong Wang , Zhaolu Kang , Youwei Liao , Yixuan Wang , Xiangyu Shi , Fengran Mo , Su Yao , Kaiyu Huang