English
Related papers

Related papers: The 2008 election: A preregistered replication ana…

200 papers

Recent research on time series foundation models has primarily focused on forecasting, leaving it unclear how generalizable their learned representations are. In this study, we examine whether frozen pre-trained forecasting models can…

Machine Learning · Computer Science 2025-10-31 Andreas Auer , Daniel Klotz , Sebastinan Böck , Sepp Hochreiter

Partisan gerrymandering is a major cause for voter disenfranchisement in United States. However, convincing US courts to adopt specific measures to quantify gerrymandering has been of limited success to date. Recently, Stephanopoulos and…

Computers and Society · Computer Science 2018-04-30 Tanima Chatterjee , Bhaskar DasGupta , Laura Palmieri , Zainab Al-Qurashi , Anastasios Sidiropoulos

When does a machine learning model predict the future of individuals and when does it recite patterns that predate the individuals? In this work, we propose a distinction between these two pathways of prediction, supported by theoretical,…

Machine Learning · Computer Science 2024-03-12 Moritz Hardt , Michael P. Kim

Solicited public opinion surveys reach a limited subpopulation of willing participants and are expensive to conduct, leading to poor time resolution and a restricted pool of expert-chosen survey topics. In this study, we demonstrate that…

Physics and Society · Physics 2016-08-09 Emily M. Cody , Andrew J. Reagan , Peter Sheridan Dodds , Christopher M. Danforth

Nationally representative surveys track public opinion, yet they ask only a limited set of questions each year, limiting its potential to capture historical changes. To fill this gap, we develop a large language model (LLM)-based framework…

Computation and Language · Computer Science 2026-05-21 Junsol Kim , Byungkyu Lee

We examine the extent of gerrymandering for the 2010 General Assembly district map of Wisconsin. We find that there is substantial variability in the election outcome depending on what maps are used. We also found robust evidence that the…

Applications · Statistics 2017-09-07 Gregory Herschlag , Robert Ravier , Jonathan C. Mattingly

Replicability analysis aims to identify the findings that replicated across independent studies that examine the same features. We provide powerful novel replicability analysis procedures for two studies for FWER and for FDR control on the…

Methodology · Statistics 2019-03-01 Marina Bogomolov , Ruth Heller

Correlation between microstructure noise and latent financial logarithmic returns is an empirically relevant phenomenon with sound theoretical justification. With few notable exceptions, all integrated variance estimators proposed in the…

Computation · Statistics 2019-05-29 Stefano Peluso , Antonietta Mira , Pietro Muliere

Large pre-trained language models are widely used in the community. These models are usually trained on unmoderated and unfiltered data from open sources like the Internet. Due to this, biases that we see in platforms online which are a…

Computation and Language · Computer Science 2023-04-17 Swapnil Sharma , Nikita Anand , Kranthi Kiran G. V. , Alind Jain

The likelihood model of high dimensional data $X_n$ can often be expressed as $p(X_n|Z_n,\theta)$, where $\theta\mathrel{\mathop:}=(\theta_k)_{k\in[K]}$ is a collection of hidden features shared across objects, indexed by $n$, and $Z_n$ is…

Machine Learning · Computer Science 2019-05-14 Aonan Zhang , John Paisley

Social media has become an emerging alternative to opinion polls for public opinion collection, while it is still posing many challenges as a passive data source, such as structurelessness, quantifiability, and representativeness. Social…

Social and Information Networks · Computer Science 2020-05-26 Zhaoya Gong , Tengteng Cai , Jean-Claude Thill , Scott Hale , Mark Graham

We study the effect of strategic behavior in iterative voting for multiple issues under uncertainty. We introduce a model synthesizing simultaneous multi-issue voting with Meir, Lev, and Rosenschein (2014)'s local dominance theory and…

Computer Science and Game Theory · Computer Science 2023-01-24 Joshua Kavner , Reshef Meir , Francesca Rossi , Lirong Xia

The memorization of training data by neural networks raises pressing concerns for privacy and security. Recent work has shown that, under certain conditions, portions of the training set can be reconstructed directly from model parameters.…

Machine Learning · Computer Science 2025-09-26 Yehonatan Refael , Guy Smorodinsky , Ofir Lindenbaum , Itay Safran

We consider a two-round election model involving $m$ voters and $n$ candidates. Each voter is endowed with a strict preference list ranking the candidates. In the first round, the candidates are partitioned into two subsets, $A$ and $B$,…

Computer Science and Game Theory · Computer Science 2026-03-17 Emilio De Santis , Antonio Di Crescenzo , Verdiana Mustaro

Reliability of machine learning evaluation -- the consistency of observed evaluation scores across replicated model training runs -- is affected by several sources of nondeterminism which can be regarded as measurement noise. Current…

Machine Learning · Computer Science 2023-10-10 Michael Hagmann , Philipp Meier , Stefan Riezler

A screening experiment attempts to identify a subset of important effects using a relatively small number of experimental runs. Given the limited run size and a large number of possible effects, penalized regression is a popular tool used…

In many randomized trials, outcomes such as essays or open-ended responses must be manually scored as a preliminary step to impact analysis, a process that is costly and limiting. Model-assisted estimation offers a way to combine surrogate…

Methodology · Statistics 2026-02-16 Reagan Mozer , Nicole E. Pashley , Luke Miratrix

We examine the necessity of interpolation in overparameterized models, that is, when achieving optimal predictive risk in machine learning problems requires (nearly) interpolating the training data. In particular, we consider simple…

Machine Learning · Statistics 2022-06-17 Chen Cheng , John Duchi , Rohith Kuditipudi

This study analyzes diverse hypotheses of electronic fraud in the Recall Referendum celebrated in Venezuela on August 15, 2004. We define fraud as the difference between the elector's intent, and the official vote tally. Our null hypothesis…

Methodology · Statistics 2012-05-18 Ricardo Hausmann , Roberto Rigobon

In Regression Discontinuity (RD) design, self-selection leads to different distributions of covariates on two sides of the policy intervention, which essentially violates the continuity of potential outcome assumption. The standard RD…

Methodology · Statistics 2019-11-22 Sida Peng , Yang Ning