English
Related papers

Related papers: Biothreat Benchmark Generation Framework for Evalu…

200 papers

The advent of Large Language Models (LLMs) has brought an unprecedented surge in machine-generated text (MGT) across diverse channels. This raises legitimate concerns about its potential misuse and societal implications. The need to…

Frontier Large Language Models (LLMs) pose unprecedented dual-use risks through the potential proliferation of chemical, biological, radiological, and nuclear (CBRN) weapons knowledge. We present the first comprehensive evaluation of 10…

Cryptography and Security · Computer Science 2025-10-27 Divyanshu Kumar , Nitin Aravind Birur , Tanay Baswa , Sahil Agarwal , Prashanth Harshangi

Assessing the capacity of Large Language Models (LLMs) to plan and reason within the constraints of interactive environments is crucial for developing capable AI agents. We introduce $\textbf{LLM-BabyBench}$, a new benchmark suite designed…

Artificial Intelligence · Computer Science 2025-05-20 Omar Choukrani , Idriss Malek , Daniil Orel , Zhuohan Xie , Zangir Iklassov , Martin Takáč , Salem Lahlou

Recent advances in the capacity of large language models to generate human-like text have resulted in their increased adoption in user-facing settings. In parallel, these improvements have prompted a heated discourse around the risks of…

Computation and Language · Computer Science 2023-02-23 Sachin Kumar , Vidhisha Balachandran , Lucille Njoo , Antonios Anastasopoulos , Yulia Tsvetkov

Large language models (LLMs) have become increasingly integrated with various applications. To ensure that LLMs do not generate unsafe responses, they are aligned with safeguards that specify what content is restricted. However, such…

Computation and Language · Computer Science 2024-05-08 Hongyu Cai , Arjun Arunasalam , Leo Y. Lin , Antonio Bianchi , Z. Berkay Celik

Retrieval-Augmented Generation allows to enhance Large Language Models with external knowledge. In response to the recent popularity of generative LLMs, many RAG approaches have been proposed, which involve an intricate number of different…

Computation and Language · Computer Science 2024-07-02 David Rau , Hervé Déjean , Nadezhda Chirkova , Thibault Formal , Shuai Wang , Vassilina Nikoulina , Stéphane Clinchant

Generative AI agents, software systems powered by Large Language Models (LLMs), are emerging as a promising approach to automate cybersecurity tasks. Among the others, penetration testing is a challenging field due to the task complexity…

Cryptography and Security · Computer Science 2024-10-29 Luca Gioacchini , Marco Mellia , Idilio Drago , Alexander Delsanto , Giuseppe Siracusano , Roberto Bifulco

Large language models (LLMs) have emerged as powerful tools for generating domain-specific multiple-choice questions (MCQs), offering efficiency gains for certification boards but raising new concerns about examination security. This study…

Computers and Society · Computer Science 2026-01-01 Ting Wang , Caroline Prendergast , Susan Lottridge

Retrieval-Augmented Generation (RAG) enhances large language models (LLMs) by integrating external knowledge sources, enabling more accurate and contextually relevant responses tailored to user queries. These systems, however, remain…

Computation and Language · Computer Science 2025-05-26 Huichi Zhou , Kin-Hei Lee , Zhonghao Zhan , Yue Chen , Zhenhao Li , Zhaoyang Wang , Hamed Haddadi , Emine Yilmaz

The rapid advancement and deployment of AI systems have created an urgent need for standard safety-evaluation frameworks. This paper introduces AILuminate v1.0, the first comprehensive industry-standard benchmark for assessing AI-product…

Computers and Society · Computer Science 2025-04-22 Shaona Ghosh , Heather Frase , Adina Williams , Sarah Luger , Paul Röttger , Fazl Barez , Sean McGregor , Kenneth Fricklas , Mala Kumar , Quentin Feuillade--Montixi , Kurt Bollacker , Felix Friedrich , Ryan Tsang , Bertie Vidgen , Alicia Parrish , Chris Knotz , Eleonora Presani , Jonathan Bennion , Marisa Ferrara Boston , Mike Kuniavsky , Wiebke Hutiri , James Ezick , Malek Ben Salem , Rajat Sahay , Sujata Goswami , Usman Gohar , Ben Huang , Supheakmungkol Sarin , Elie Alhajjar , Canyu Chen , Roman Eng , Kashyap Ramanandula Manjusha , Virendra Mehta , Eileen Long , Murali Emani , Natan Vidra , Benjamin Rukundo , Abolfazl Shahbazi , Kongtao Chen , Rajat Ghosh , Vithursan Thangarasa , Pierre Peigné , Abhinav Singh , Max Bartolo , Satyapriya Krishna , Mubashara Akhtar , Rafael Gold , Cody Coleman , Luis Oala , Vassil Tashev , Joseph Marvin Imperial , Amy Russ , Sasidhar Kunapuli , Nicolas Miailhe , Julien Delaunay , Bhaktipriya Radharapu , Rajat Shinde , Tuesday , Debojyoti Dutta , Declan Grabb , Ananya Gangavarapu , Saurav Sahay , Agasthya Gangavarapu , Patrick Schramowski , Stephen Singam , Tom David , Xudong Han , Priyanka Mary Mammen , Tarunima Prabhakar , Venelin Kovatchev , Rebecca Weiss , Ahmed Ahmed , Kelvin N. Manyeki , Sandeep Madireddy , Foutse Khomh , Fedor Zhdanov , Joachim Baumann , Nina Vasan , Xianjun Yang , Carlos Mougn , Jibin Rajan Varghese , Hussain Chinoy , Seshakrishna Jitendar , Manil Maskey , Claire V. Hardgrove , Tianhao Li , Aakash Gupta , Emil Joswin , Yifan Mai , Shachi H Kumar , Cigdem Patlak , Kevin Lu , Vincent Alessi , Sree Bhargavi Balija , Chenhe Gu , Robert Sullivan , James Gealy , Matt Lavrisa , James Goel , Peter Mattson , Percy Liang , Joaquin Vanschoren

Retrieval-Augmented Generation (RAG) enhances large language models (LLMs) by incorporating external knowledge, but its openness introduces vulnerabilities that can be exploited by poisoning attacks. Existing poisoning methods for RAG…

Cryptography and Security · Computer Science 2025-05-27 Chunyang Li , Junwei Zhang , Anda Cheng , Zhuo Ma , Xinghua Li , Jianfeng Ma

Cybersecurity spans multiple interconnected domains, complicating the development of meaningful, labor-relevant benchmarks. Existing benchmarks assess isolated skills rather than integrated performance. We find that pre-trained knowledge of…

Large language models (LLMs) can enhance factuality via retrieval-augmented generation (RAG), but applying RAG to every query is unnecessary when the model-only answer is reliable. This motivates cascaded RAG: each query is first handled by…

Computation and Language · Computer Science 2026-05-20 Zijun Jia , Yuanchang Ye , Sen Jia , Yiyao Qian , Haoning Wang , Baojie Chen , Diyin Tang , Jinsong Yu , Zhiyuan Wang

The era of large language models (LLM) raises questions not only about how to train models, but also about how to evaluate them. Despite numerous existing benchmarks, insufficient attention is often given to creating assessments that test…

We introduce BikeBench, an engineering design benchmark for evaluating generative models on problems with multiple real-world objectives and constraints. As generative AI's reach continues to grow, evaluating its capability to understand…

Computational Engineering, Finance, and Science · Computer Science 2025-10-28 Lyle Regenwetter , Yazan Abu Obaideh , Fabien Chiotti , Ioanna Lykourentzou , Faez Ahmed

Large Language Models (LLMs) are widely deployed in diverse real-world settings, yet remain vulnerable to jailbreaking, where prompt-based attacks bypass safety filters. We present THREAT (Targeted Harmful generation via Reframing and…

Cryptography and Security · Computer Science 2026-05-22 Shahnewaz Karim Sakib , Swati Kar , Anindya Bijoy Das

As software systems grow more complex, automated testing has become essential to ensuring reliability and performance. Traditional methods for boundary value test input generation can be time-consuming and may struggle to address all…

Software Engineering · Computer Science 2025-01-27 Xiujing Guo , Chen Li , Tatsuhiro Tsuchiya

Large Language Models (LLMs) are increasingly integrated into critical decision-making pipelines, a trend that raises the demand for robust and automated data analysis. Current approaches to dataset risk analysis are limited to manual…

Artificial Intelligence · Computer Science 2026-05-28 Panteleimon Rodis

Risk thresholds provide a measure of the level of risk exposure that a society or individual is willing to withstand, ultimately shaping how we determine the safety of technological systems. Against the backdrop of the Cold War, the first…

Computers and Society · Computer Science 2025-04-22 Heidy Khlaaf , Sarah Myers West

Advanced AI systems offer substantial benefits but also introduce risks. In 2025, AI-enabled cyber offense has emerged as a concrete example. This technical report applies a quantitative risk modeling methodology (described in full in a…

‹ Prev 1 4 5 6 7 8 10 Next ›