相关论文: Comparing Apples and Oranges: Two Examples of the …
We investigate the generalizability of learned binary relations: functions that map pairs of instances to a logical indicator. This problem has application in numerous areas of machine learning, such as ranking, entity resolution and link…
The angular measure on the unit sphere characterizes the first-order dependence structure of the components of a random vector in extreme regions and is defined in terms of standardized margins. Its statistical recovery is an important step…
Researchers currently use a number of approaches to predict and substantiate information-computation gaps in high-dimensional statistical estimation problems. A prominent approach is to characterize the limits of restricted models of…
Machine learning models play a key role for service providers looking to gain market share in consumer markets. However, traditional learning approaches do not take into account the existence of additional providers, who compete with each…
Online controlled experiments, or A/B tests, are large-scale randomized trials in digital environments. This paper investigates the estimands of the difference-in-means estimator in these experiments, focusing on scenarios with repeated…
The standard central limit theorem with a Gaussian attractor for the sum of independent random variables may lose its validity in presence of strong correlations between the added random contributions. Here, we study this problem for…
Estimating the dependences between random variables, and ranking them accordingly, is a prevalent problem in machine learning. Pursuing frequentist and information-theoretic approaches, we first show that the p-value and the mutual…
Binary classification is widely used in ML production systems. Monitoring classifiers in a constrained event space is well known. However, real world production systems often lack the ground truth these methods require. Privacy concerns may…
In statistical problems, a set of parameterized probability distributions is used to estimate the true probability distribution. If Fisher information matrix at the true distribution is singular, then it has been left unknown what we can…
The personalization of our news consumption on social media has a tendency to reinforce our pre-existing beliefs instead of balancing our opinions. This finding is a concern for the health of our democracies which rely on an access to…
Large language models (LLMs) are increasingly used to meet user information needs, but their effectiveness in dealing with user queries that contain various types of ambiguity remains unknown, ultimately risking user trust and satisfaction.…
I study the measurement of scientists' influence using bibliographic data. The main result is an axiomatic characterization of the family of citation-counting indices, a broad class of influence measures which includes the renowned h-index.…
The quantum Cram\'er-Rao bound sets a fundamental limit on the accuracy of unbiased parameter estimation in quantum systems, relating the uncertainty in determining a parameter to the inverse of the quantum Fisher information. We…
How does targeted advertising influence electoral outcomes? This paper presents a one-dimensional spatial model of voting in which a privately informed challenger persuades voters to support him over the status quo. I show that targeted…
Political advertising on social media has become a central element in election campaigns. However, granular information about political advertising on social media was previously unavailable, thus raising concerns regarding fairness,…
Repeated sampling is a standard way to spend test-time compute, but its benefit is controlled by the latent distribution of correctness across examples, not by one-call accuracy alone. We study the binary correctness layer of repeated LLM…
In the Hegselmann-Krause model, an agent updates its opinion by averaging with others whose opinions differ by at most a given confidence threshold. With agents' initial opinions uniformly distributed on the unit interval, we provide a…
Randomized experiments with treatment and control groups are an important tool to measure the impacts of interventions. However, in experimental settings with one-sided noncompliance extant empirical approaches may not produce the estimands…
We consider two variants of a quantum-statistical generalization of the Cramer-Rao inequality that establishes an invariant lower bound on the mean square error of a generalized quantum measurement. The proposed complex variant of this…
In a recent SIGMOD paper titled "Debunking the Myths of Influence Maximization: An In-Depth Benchmarking Study", Arora et al. [1] undertake a performance benchmarking study of several well-known algorithms for influence maximization. In the…