Related papers: A Weakness Measure for GR(1) Formulae
While pre-trained language models (LMs) have brought great improvements in many NLP tasks, there is increasing attention to explore capabilities of LMs and interpret their predictions. However, existing works usually focus only on a certain…
Sustainability reports are critical for ESG assessment, yet greenwashing and vague claims often undermine their reliability. Existing NLP models lack robustness to these practices, typically relying on surface-level patterns that generalize…
The recent successful paradigm of solving logical reasoning problems with tool-augmented large language models (LLMs) leverages translation of natural language (NL) statements into First-Order Logic~(FOL) and external theorem provers.…
Tests of general relativity (GR) with gravitational waves (GWs) introduce additional deviation parameters in the waveform model. The enlarged parameter space makes inference computationally costly, which has so far limited systematic,…
Quantum measurement is one of the most fascinating and discussed phenomena in quantum physics, due to the impact on the system of the measurement action and the resulting interpretation issues. Scholars proposed weak measurements to amplify…
Objective metrics for emotional expressiveness are vital for speech generation, particularly in expressive synthesis and voice conversion requiring emotional prosody transfer. To quantify this, the field widely relies on emotion similarity…
We initiate a formal study of reproducibility in optimization. We define a quantitative measure of reproducibility of optimization procedures in the face of noisy or error-prone operations such as inexact or stochastic gradient computations…
Binary 0-1 measurement matrices, especially those from coding theory, were introduced to compressed sensing (CS) recently. Good measurement matrices with preferred properties, e.g., the restricted isometry property (RIP) and nullspace…
We prove near-optimal trade-offs for quantifier depth versus number of variables in first-order logic by exhibiting pairs of $n$-element structures that can be distinguished by a $k$-variable first-order sentence but where every such…
We introduce a notion of weak convergence in arbitrary metric spaces. Metric functionals are key in our analysis: weak convergence of sequences in a given metric space is tested against all the metric functionals defined on said space. When…
In the paper we introduce a weak set theory $\mathsf{H}_{<\omega}$ . A formalization of arithmetic on finite von Neumann ordinals gives an embedding of arithmetical language into this theory. We show that $\mathsf{H}_{<\omega}$ proves a…
In this paper, we investigate the sample complexity of policy evaluation in infinite-horizon offline reinforcement learning (also known as the off-policy evaluation problem) with linear function approximation. We identify a hard regime…
Post-training has become central to improving reasoning and alignment in large language models, where critic-free models enable scalable learning from model-generated outputs but lack principled mechanisms to distinguish informative from…
We study weighted norm inequalities of $(1,q)$- type for $0<q<1$, $\Vert \mathbf{G} \nu \Vert_{L^q(\Omega, d \sigma)} \le C \, \Vert \nu \Vert, \quad \text{for all positive measures $\nu$ in $\Omega$},$ along with their weak-type…
The quantum nature of neutrino oscillations would be reflected in the mismatch between the neutrino survival probabilities with and without an intermediate observation. We propose this ``quantum mismatch'' as a measure of quantumness in…
The increasing sensitivity of current and upcoming gravitational-wave (GW) detectors poses stringent requirements on the accuracy of the GW models used for data analysis. If these requirements are not met, systematic errors could dominate…
We present a general theory to quantify the uncertainty from imposing structural assumptions on the second-order structure of nonstationary Hilbert space-valued processes, which can be measured via functionals of time-dependent spectral…
We present a method to quantify a system's resilience capacity, i.e., the set of degradation magnitudes for which all functional requirements remain satisfied. These requirements come from human stakeholders (e.g., operators, planners) who…
The robustness of classifiers has become a question of paramount importance in the past few years. Indeed, it has been shown that state-of-the-art deep learning architectures can easily be fooled with imperceptible changes to their inputs.…
Quantifying robustness in a single measure for the purposes of model selection, development of adversarial training methods, and anticipating trends has so far been elusive. The simplest metric to consider is the number of trainable…