English
Related papers

Related papers: RedacBench: Can AI Erase Your Secrets?

200 papers

Evaluating Large Language Models (LLMs) on repository-level feature implementation is a critical frontier in software engineering. However, establishing a benchmark that faithfully mirrors realistic development scenarios remains a…

Computation and Language · Computer Science 2026-02-19 Haorui Chen , Chengze Li , Jia Li

The increasing realism of synthetic speech, driven by advancements in text-to-speech models, raises ethical concerns regarding impersonation and disinformation. Audio watermarking offers a promising solution via embedding…

Machine Learning · Computer Science 2024-11-14 Hongbin Liu , Moyang Guo , Zhengyuan Jiang , Lun Wang , Neil Zhenqiang Gong

We present a comprehensive evaluation of large language models for multilingual readability assessment. Existing evaluation resources lack domain and language diversity, limiting the ability for cross-domain and cross-lingual analyses. This…

Computation and Language · Computer Science 2024-10-17 Tarek Naous , Michael J. Ryan , Anton Lavrouk , Mohit Chandra , Wei Xu

Data containing personal information is increasingly used to train, fine-tune, or query Large Language Models (LLMs). Text is typically scrubbed of identifying information prior to use, often with tools such as Microsoft's Presidio or…

Computation and Language · Computer Science 2026-02-16 Nataša Krčo , Zexi Yao , Matthieu Meeus , Yves-Alexandre de Montjoye

As large language models continue to advance, their application in educational contexts remains underexplored and under-optimized. In this paper, we address this gap by introducing the first diverse benchmark tailored for educational…

Computation and Language · Computer Science 2026-01-07 Bin Xu , Yu Bai , Huashan Sun , Yiguan Lin , Siming Liu , Xinyue Liang , Yaolin Li , Zhuangzhi Dong , Jingren Zhang , Yufan Deng , Xinyu Zou , Yang Gao , Heyan Huang

Real-world data analysis tasks often come with under-specified goals and unclean data. User interaction is necessary to understand and disambiguate a user's intent, and hence, essential to solving these complex tasks. Existing benchmarks…

Uncontrollable autonomous replication of language model agents poses a critical safety risk. To better understand this risk, we introduce RepliBench, a suite of evaluations designed to measure autonomous replication capabilities. RepliBench…

Evaluating language models fairly is increasingly difficult as static benchmarks risk contamination by training data, obscuring whether models truly reason or recall. We introduce BeyondBench, an evaluation framework using algorithmic…

Computation and Language · Computer Science 2026-03-06 Gaurav Srivastava , Aafiya Hussain , Zhenyu Bi , Swastik Roy , Priya Pitre , Meng Lu , Morteza Ziyadi , Xuan Wang

Language model agents are increasingly used to automate scientific research, yet evaluating their scientific contributions remains a challenge. A key mechanism to obtain such insights is through ablation experiments. To this end, we…

Computation and Language · Computer Science 2026-02-03 Talor Abramovich , Gal Chechik

As large language models (LLMs) rapidly advance and integrate into daily life, the privacy risks they pose are attracting increasing attention. We focus on a specific privacy risk where LLMs may help identify the authorship of anonymous…

Computation and Language · Computer Science 2024-11-21 Zichen Wen , Dadi Guo , Huishuai Zhang

Differentially private (DP) synthetic data generation is a promising technique for utilizing private datasets that otherwise cannot be exposed for model training or other analytics. While much research literature has focused on generating…

Computation and Language · Computer Science 2025-09-16 Shuaiqi Wang , Vikas Raunak , Arturs Backurs , Victor Reis , Pei Zhou , Sihao Chen , Longqi Yang , Zinan Lin , Sergey Yekhanin , Giulia Fanti

Recent advances in Retrieval-Augmented Generation (RAG) have enabled large language models (LLMs) to ground outputs in clinical evidence. However, connecting LLMs with external databases introduces the risk of contextual leakage: a subtle…

Computation and Language · Computer Science 2026-03-17 Shaowei Guan , Yu Zhai , Hin Chi Kwok , Jiawei Du , Xinyu Feng , Jing Li , Harry Qin , Vivian Hui

Frontier model progress is often measured by academic benchmarks, which offer a limited view of performance in real-world professional contexts. Existing evaluations often fail to assess open-ended, economically consequential tasks in…

Despite the remarkable advances of Large Language Models (LLMs) across diverse cognitive tasks, the rapid enhancement of these capabilities also introduces emergent deceptive behaviors that may induce severe risks in high-stakes…

Computation and Language · Computer Science 2025-11-18 Yao Huang , Yitong Sun , Yichi Zhang , Ruochen Zhang , Yinpeng Dong , Xingxing Wei

We introduce AuditBench, an alignment auditing benchmark. AuditBench consists of 56 language models with implanted hidden behaviors. Each model has one of 14 concerning behaviors--such as sycophantic deference, opposition to AI regulation,…

Computation and Language · Computer Science 2026-03-11 Abhay Sheshadri , Aidan Ewart , Kai Fronsdal , Isha Gupta , Samuel R. Bowman , Sara Price , Samuel Marks , Rowan Wang

Text segmentation is a prerequisite in many real-world text-related tasks, e.g., text style transfer, and scene text removal. However, facing the lack of high-quality datasets and dedicated investigations, this critical prerequisite has…

Computer Vision and Pattern Recognition · Computer Science 2020-12-01 Xingqian Xu , Zhifei Zhang , Zhaowen Wang , Brian Price , Zhonghao Wang , Humphrey Shi

We present HealthBench, an open-source benchmark measuring the performance and safety of large language models in healthcare. HealthBench consists of 5,000 multi-turn conversations between a model and an individual user or healthcare…

We introduce SpreadsheetBench, a challenging spreadsheet manipulation benchmark exclusively derived from real-world scenarios, designed to immerse current large language models (LLMs) in the actual workflow of spreadsheet users. Unlike…

Computation and Language · Computer Science 2024-10-18 Zeyao Ma , Bohan Zhang , Jing Zhang , Jifan Yu , Xiaokang Zhang , Xiaohan Zhang , Sijia Luo , Xi Wang , Jie Tang

Recent advances in multi-modal generative models have driven substantial improvements in image editing. However, current generative models still struggle with handling diverse and complex image editing tasks that require implicit reasoning,…

Computer Vision and Pattern Recognition · Computer Science 2025-11-25 Feng Han , Yibin Wang , Chenglin Li , Zheming Liang , Dianyi Wang , Yang Jiao , Zhipeng Wei , Chao Gong , Cheng Jin , Jingjing Chen , Jiaqi Wang

The creation of benchmarks to evaluate the safety of Large Language Models is one of the key activities within the trusted AI community. These benchmarks allow models to be compared for different aspects of safety such as toxicity, bias,…

Artificial Intelligence · Computer Science 2025-06-23 Lina Berrayana , Sean Rooney , Luis Garcés-Erice , Ioana Giurgiu