English
Related papers

Related papers: When Alpha Disappears: A One-Switch Benchmark for …

200 papers

The performance of large language models (LLMs) continues to improve, as reflected in rising scores on standard benchmarks. However, the lack of transparency around training data raises concerns about potential overlap with evaluation sets…

Computation and Language · Computer Science 2025-06-02 Naila Shafirni Hidayat , Muhammad Dehan Al Kautsar , Alfan Farizki Wicaksono , Fajri Koto

Privacy regulations require the erasure of data from deep learning models. This is a significant challenge that is amplified in Federated Learning, where data remains on clients, making full retraining or coordinated updates often…

Machine Learning · Computer Science 2026-01-27 Antonio Balordi , Lorenzo Manini , Fabio Stella , Alessio Merlo

ForesightFlow is an Information Leakage Score (ILS) framework for detecting informed trading on decentralized prediction markets. For an event-resolved binary market, the score quantifies the fraction of the terminal information move priced…

Trading and Market Microstructure · Quantitative Finance 2026-05-15 Maksym Nechepurenko

Linear algebra expressions, which play a central role in countless scientific computations, are often computed via a sequence of calls to existing libraries of building blocks (such as those provided by BLAS and LAPACK). A sequence…

Performance · Computer Science 2024-08-15 Aravind Sankaran , Paolo Bientinesi

With online payment platforms being ubiquitous and important, fraud transaction detection has become the key for such platforms, to ensure user account safety and platform security. In this work, we present a novel method for detecting…

Machine Learning · Computer Science 2020-03-30 Longfei Li , Ziqi Liu , Chaochao Chen , Ya-Lin Zhang , Jun Zhou , Xiaolong Li

Building natural language inference (NLI) benchmarks that are both challenging for modern techniques, and free from shortcut biases is difficult. Chief among these biases is "single sentence label leakage," where annotator-introduced…

Computation and Language · Computer Science 2023-02-14 Michael Saxon , Xinyi Wang , Wenda Xu , William Yang Wang

Last year we argued that if slow-roll inflation followed the decay of a false vacuum in a large landscape, the steepening of the scalar potential between the inflationary plateau and the barrier generically leads to a potentially observable…

Cosmology and Nongalactic Astrophysics · Physics 2015-06-19 Raphael Bousso , Daniel Harlow , Leonardo Senatore

Membership Inference Attacks exploit the vulnerabilities of exposing models trained on customer data to queries by an adversary. In a recently proposed implementation of an auditing tool for measuring privacy leakage from sensitive…

Machine Learning · Computer Science 2020-09-21 Abhinav Aggarwal , Zekun Xu , Oluwaseyi Feyisetan , Nathanael Teissier

This paper introduces a novel sparse latent factor modeling framework using sparse asymptotic Principal Component Analysis (APCA) to analyze the co-movements of high-dimensional panel data over time. Unlike existing methods based on sparse…

Methodology · Statistics 2025-08-08 Zhaoxing Gao

While traditional equity factor investing relies heavily on slow-moving fundamental accounting metrics, these models frequently suffer from factor crowding and miss real-time, sentiment-driven market dislocations. This study explores how…

Statistical Finance · Quantitative Finance 2026-05-22 Jin Du , Alexander Walter , Maxim Ulrich

The rise of digital payments has accelerated the need for intelligent and scalable systems to detect fraud. This research presents an end-to-end, feature-rich machine learning framework for detecting credit card transaction anomalies and…

In this work we propose a statistical approach to handling sources of theoretical uncertainty in string theory models of inflation. By viewing a model of inflation as a probabilistic graph, we show that there is an inevitable information…

High Energy Physics - Theory · Physics 2019-06-05 Mafalda Dias , Jonathan Frazer , Alexander Westphal

Benchmark-based evaluation is the de facto standard for comparing large language models (LLMs). However, its reliability is increasingly threatened by test set contamination, where test samples or their close variants leak into training…

Computation and Language · Computer Science 2026-01-28 Jianzhe Chai , Yu Zhe , Jun Sakuma

We consider testing zero pricing errors in high-dimensional linear factor pricing models. Existing methods are mainly based on either an $L_2$ statistic, which is effective under dense alternatives, or an $L_\infty$ statistic, which is…

Methodology · Statistics 2026-04-01 Ping Zhao , Huifang Ma , Long Feng

In the field of fraud detection, the availability of comprehensive and privacy-compliant datasets is crucial for advancing machine learning research and developing effective anti-fraud systems. Traditional datasets often focus on…

Machine Learning · Computer Science 2024-04-24 Phoebe Jing , Yijing Gao , Xianlong Zeng

Large Language Models (LLMs) are trained on massive web-crawled corpora. This poses risks of leakage, including personal information, copyrighted texts, and benchmark datasets. Such leakage leads to undermining human trust in AI due to…

Computation and Language · Computer Science 2024-03-26 Masahiro Kaneko , Timothy Baldwin

We consider the weakly supervised binary classification problem where the labels are randomly flipped with probability $1- {\alpha}$. Although there exist numerous algorithms for this problem, it remains theoretically unexplored how the…

Machine Learning · Computer Science 2019-07-16 Xinyang Yi , Zhaoran Wang , Zhuoran Yang , Constantine Caramanis , Han Liu

We introduce a family of information leakage measures called maximal $(\alpha,\beta)$-leakage (M$\alpha$beL), parameterized by real numbers $\alpha$ and $\beta$ greater than or equal to 1. The measure is formalized via an operational…

Information Theory · Computer Science 2024-04-08 Atefeh Gilani , Gowtham R. Kurri , Oliver Kosut , Lalitha Sankar

Neural networks applied to financial time series operate in a regime of underspecification, where model predictors achieve indistinguishable out-of-sample error. Using large-scale volatility forecasting for S$\&$P 500 stocks, we show that…

Machine Learning · Computer Science 2026-03-04 Federico Vittorio Cortesi , Giuseppe Iannone , Giulia Crippa , Tomaso Poggio , Pierfrancesco Beneventano

While cryptographic algorithms such as the ubiquitous Advanced Encryption Standard (AES) are secure, *physical implementations* of these algorithms in hardware inevitably 'leak' sensitive data such as cryptographic keys. A particularly…

Machine Learning · Computer Science 2026-03-26 Jimmy Gammell , Anand Raghunathan , Abolfazl Hashemi , Kaushik Roy