中文
相关论文

相关论文: GLiNER Guard: Unified Encoder Family for Productio…

200 篇论文

Ensuring safe, policy-compliant outputs from large language models requires real-time content moderation that can scale across multiple safety dimensions. However, state-of-the-art guardrail models rely on autoregressive decoders with…

计算与语言 · 计算机科学 2026-05-11 Urchade Zaratiana , Mary Newhauser , George Hurn-Maloney , Ash Lewis

Large Language Models (LLMs) hold promise for advancing legal practice by automating complex tasks and improving access to justice. However, their adoption is limited by concerns over client confidentiality, especially when lawyers include…

计算与语言 · 计算机科学 2025-01-22 M. Mikail Demir , Hakan T. Otal , M. Abdullah Canbaz

As large language models (LLMs) advance, ensuring AI safety and alignment is paramount. One popular approach is prompt guards, lightweight mechanisms designed to filter malicious queries while being easy to implement and update. In this…

机器学习 · 计算机科学 2025-10-08 Jaiden Fairoze , Sanjam Garg , Keewoo Lee , Mingyuan Wang

We introduce WildGuard -- an open, light-weight moderation tool for LLM safety that achieves three goals: (1) identifying malicious intent in user prompts, (2) detecting safety risks of model responses, and (3) determining model refusal…

计算与语言 · 计算机科学 2024-12-11 Seungju Han , Kavel Rao , Allyson Ettinger , Liwei Jiang , Bill Yuchen Lin , Nathan Lambert , Yejin Choi , Nouha Dziri

Production LLM systems often rely on separate models for safety and other classification-heavy steps, increasing latency, VRAM footprint, and operational complexity. We instead reuse computation already paid for by the serving LLM: we train…

计算与语言 · 计算机科学 2026-04-28 Gonzalo Ariel Meyoyan , Luciano Del Corro

Reliable detection of personally identifiable information (PII) is increasingly important across modern data-processing systems, yet the task remains difficult: PII spans are heterogeneous, locale-dependent, context-sensitive, and often…

计算与语言 · 计算机科学 2026-05-12 Urchade Zaratiana , Ash Lewis , George Hurn-Maloney

Large language models (LLMs) are increasingly deployed in high-stakes domains, yet a unified treatment of their overlapping safety challenges remains lacking. We present SafeLM, a framework that jointly addresses four pillars of LLM safety:…

密码学与安全 · 计算机科学 2026-04-21 Noor Islam S. Mohammad , Uluğ Bayazıt

Omni-modal Large Language Models (OLLMs) that process text, images, videos, and audio introduce new challenges for safety and value guardrails in human-AI interaction. Prior guardrail research largely targets unimodal settings and typically…

人工智能 · 计算机科学 2025-12-03 Boyu Zhu , Xiaofei Wen , Wenjie Jacky Mo , Tinghui Zhu , Yanan Xie , Peng Qi , Muhao Chen

The adoption of Large Language Models (LLMs) has revolutionized AI applications but poses significant challenges in safeguarding user privacy. Ensuring compliance with privacy regulations such as GDPR and CCPA while addressing nuanced…

密码学与安全 · 计算机科学 2025-01-23 Shubhi Asthana , Bing Zhang , Ruchi Mahindru , Chad DeLuca , Anna Lisa Gentile , Sandeep Gopisetty

As Large Language Models (LLMs) grow increasingly powerful, ensuring their safety and alignment with human values remains a critical challenge. Ideally, LLMs should provide informative responses while avoiding the disclosure of harmful or…

计算与语言 · 计算机科学 2024-10-04 Lingrui Mei , Shenghua Liu , Yiwei Wang , Baolong Bi , Ruibin Yuan , Xueqi Cheng

In many practical LLM deployments, a single guardrail is used for both prompt and response moderation. Prompt moderation operates on fully observed text, whereas streaming response moderation requires safety decisions to be made over…

计算与语言 · 计算机科学 2026-04-07 Pride Kavumba , Koki Wataoka , Huy H. Nguyen , Jiaxuan Li , Masaya Ohagi

The growing adoption of Large Language Models (LLMs) has influenced the development of Small Language Models (SLMs) for on-device deployment across smartphones and edge devices, offering enhanced privacy, reduced latency, server-free…

密码学与安全 · 计算机科学 2026-03-04 Kalyan Nakka , Jimmy Dani , Ausmit Mondal , Nitesh Saxena

Ensuring the safety of LLM-generated content is essential for real-world deployment. Most existing guardrail models formulate moderation as a fixed binary classification task, implicitly assuming a fixed definition of harmfulness. In…

机器学习 · 计算机科学 2026-04-16 Zhihao Ding , Jinming Li , Ze Lu , Jieming Shi

We present SGuard-v1, a lightweight safety guardrail for Large Language Models (LLMs), which comprises two specialized models to detect harmful content and screen adversarial prompts in human-AI conversational settings. The first component,…

计算与语言 · 计算机科学 2025-11-18 JoonHo Lee , HyeonMin Cho , Jaewoong Yun , Hyunjae Lee , JunKyu Lee , Juree Seok

Large Language Models (LLMs) are powerful tools for answering user queries, yet they remain highly vulnerable to jailbreak attacks. Existing guardrail methods typically rely on internal features or textual responses to detect malicious…

密码学与安全 · 计算机科学 2026-05-29 Zikai Zhang , Rui Hu , Olivera Kotevska , Jiahao Xu

Automated masking of Personally Identifiable Information (PII) is critical for privacy-preserving conversational systems. While current frontier large language models demonstrate strong PII masking capabilities, concerns about data handling…

计算与语言 · 计算机科学 2025-12-23 Prabigya Acharya , Liza Shrestha

Multi-agent LLM systems introduce a security risk in which sensitive information accessed by one agent can propagate through shared context and reappear in downstream outputs, even without explicit adversarial intent. We formalise this…

人工智能 · 计算机科学 2026-05-12 Riya Tapwal , Abhishek Kumar , Carsten Maple

The proliferation of AI technology gives rise to a variety of security threats, which significantly compromise the confidentiality and integrity of AI models and applications. Existing software-based solutions mainly target one specific…

密码学与安全 · 计算机科学 2023-11-29 Xiaobei Yan , Han Qiu , Tianwei Zhang

Guard models are widely used to detect harmful content in user prompts and LLM responses. However, state-of-the-art guard models rely solely on terminal-layer representations and overlook the rich safety-relevant features distributed across…

人工智能 · 计算机科学 2026-04-21 Difan Jiao , Yilun Liu , Ye Yuan , Zhenwei Tang , Linfeng Du , Haolun Wu , Ashton Anderson

We present ShieldGemma, a comprehensive suite of LLM-based safety content moderation models built upon Gemma2. These models provide robust, state-of-the-art predictions of safety risks across key harm types (sexually explicit, dangerous…

‹ 上一页 1 2 3 10 下一页 ›