English
Related papers

Related papers: Data Leakage and Redundancy in the LIT-PCBA Benchm…

200 papers

With the widespread application of artificial intelligence technologies in face recognition and other fields, data privacy security issues have received extensive attention, especially the \textit{right to be forgotten} emphasized by…

Cryptography and Security · Computer Science 2026-04-10 Weidong Zheng , Kongyang Chen , Yao Huang , Yuanwei Guo , Yatie Xiao

As AI agents become integral to enterprise workflows, their reliance on shared tool libraries and pre-trained components creates significant supply chain vulnerabilities. While previous work has demonstrated behavioral backdoor detection…

Cryptography and Security · Computer Science 2025-11-26 Arun Chowdary Sanna

Analog in-memory computing (AIMC) is a promising compute paradigm to improve speed and power efficiency of neural network inference beyond the limits of conventional von Neumann-based architectures. However, AIMC introduces fundamental…

Twenty-eight within-subject counterfactual experiments across 2,047 tabular datasets, plus a boundary experiment on 129 temporal datasets, measuring the severity of four data leakage classes in machine learning. Class I (estimation -…

Machine Learning · Computer Science 2026-04-07 Simon Roth

Generative large language models are crucial in natural language processing, but they are vulnerable to backdoor attacks, where subtle triggers compromise their behavior. Although backdoor attacks against LLMs are constantly emerging,…

Cryptography and Security · Computer Science 2025-02-27 Xuxu Liu , Siyuan Liang , Mengya Han , Yong Luo , Aishan Liu , Xiantao Cai , Zheng He , Dacheng Tao

Analog optical computers promise large efficiency gains for machine learning inference, yet no demonstration has moved beyond small-scale image benchmarks. We benchmark the analog optical computer (AOC) digital twin on mortgage approval…

Machine Learning · Computer Science 2026-04-16 Sofia Berloff , Pavel Koptev , Konstantin Malkov

Machine unlearning has the potential to improve the safety of large language models (LLMs) by removing sensitive or harmful information post hoc. A key challenge in unlearning involves balancing between forget quality (effectively…

Machine Learning · Computer Science 2025-06-23 Shengyuan Hu , Neil Kale , Pratiksha Thaker , Yiwei Fu , Steven Wu , Virginia Smith

Graph topology is a fundamental determinant of memory leakage in multi-agent LLM systems, yet its effects remain poorly quantified. We introduce MAMA (Multi-Agent Memory Attack), a framework that measures how network structure shapes…

Cryptography and Security · Computer Science 2026-01-13 Jinbo Liu , Defu Cao , Yifei Wei , Tianyao Su , Yuan Liang , Yushun Dong , Yan Liu , Yue Zhao , Xiyang Hu

Readout of superconducting qubits faces a trade-off between measurement speed and unwanted back-action on the qubit caused by the readout drive, such as $T_1$ degradation and leakage out of the computational subspace. The readout is…

Quantum Physics · Physics 2025-07-08 S. Hazra , W. Dai , T. Connolly , P. D. Kurilovich , Z. Wang , L. Frunzio , M. H. Devoret

The growing capabilities of Large Language Models (LLMs) show significant potential to enhance healthcare by assisting medical researchers and physicians. However, their reliance on static training data is a major risk when medical…

Computation and Language · Computer Science 2025-09-05 Juraj Vladika , Mahdi Dhaini , Florian Matthes

Molecular representation learning is pivotal for various molecular property prediction tasks related to drug discovery. Robust and accurate benchmarks are essential for refining and validating current methods. Existing molecular property…

Chemical Physics · Physics 2024-06-27 Shikun Feng , Jiaxin Zheng , Yinjun Jia , Yanwen Huang , Fengfeng Zhou , Wei-Ying Ma , Yanyan Lan

The recent progress in text-based audio retrieval was largely propelled by the release of suitable datasets. Since the manual creation of such datasets is a laborious task, obtaining data from online resources can be a cheap solution to…

Sound · Computer Science 2023-08-29 Benno Weck , Xavier Serra

The integration of large language models (LLMs) into cyber security applications presents both opportunities and critical safety risks. We introduce CyberLLMInstruct, a dataset of 54,928 pseudo-malicious instruction-response pairs spanning…

Cryptography and Security · Computer Science 2025-09-18 Adel ElZemity , Budi Arief , Shujun Li

Recently, various techniques (e.g., fuzzing) have been developed for vulnerability detection. To evaluate those techniques, the community has been developing benchmarks of artificial vulnerabilities because of a shortage of ground-truth.…

Cryptography and Security · Computer Science 2020-03-24 Sijia Geng , Yuekang Li , Yunlan Du , Jun Xu , Yang Liu , Bing Mao

Meta-Continual Learning (Meta-CL) enables models to learn new classes from limited labelled samples, making it promising for IoT applications where manual labelling is costly. However, existing studies focus on accuracy while ignoring…

Machine Learning · Computer Science 2026-01-27 Sijia Li , Young D. Kwon , Lik-Hang Lee , Pan Hui

Lithography modeling is a crucial problem in chip design to ensure a chip design mask is manufacturable. It requires rigorous simulations of optical and chemical models that are computationally expensive. Recent developments in machine…

We study bitstring measurements from the publicly available Aquila Rydberg-atom platform using a two-leg ladder that encodes a truncated lattice gauge model as a practical benchmark that can be directly implemented and simulated on current…

LLM-based automated program repair (APR) techniques have shown promising results in reducing debugging costs. However, prior results can be affected by data leakage: large language models (LLMs) may memorize bug fixes when evaluation…

Software Engineering · Computer Science 2026-04-24 Milan De Koning , Ali Asgari , Pouria Derakhshanfar , Annibale Panichella

Machine learning potential enables molecular dynamics simulations of systems beyond the capability of classical force fields. The traditional approach to develop structural sets for training machine learning potential typically generate a…

Computational Physics · Physics 2021-09-06 Nan Xu , Chen Li , Mandi Fang , Qing Shao , Yingying Lu , Yao Shi , Yi He

Machine Learning (ML) has revolutionized various domains, offering predictive capabilities in several areas. However, with the increasing accessibility of ML tools, many practitioners, lacking deep ML expertise, adopt a "push the button"…

Machine Learning · Computer Science 2025-08-21 Andrea Apicella , Francesco Isgrò , Roberto Prevete