English
Related papers

Related papers: Protocols for Observational Studies: An Applicatio…

200 papers

Partial Least Squares (PLS) regression emerged as an alternative to ordinary least squares for addressing multicollinearity in a wide range of scientific applications. As multidimensional tensor data is becoming more widespread, tensor…

Methodology · Statistics 2024-10-11 Kwangmoon Park , Sündüz Keleş

We present simple low-level conditions for identification in regression discontinuity designs using a potential outcome framework for the manipulation of the running variable. Using this framework, we replace the existing identification…

Econometrics · Economics 2024-09-18 Takuya Ishihara , Masayuki Sawada

Using offline observational data for policy evaluation and learning allows decision-makers to evaluate and learn a policy that connects characteristics and interventions. Most existing literature has focused on either discrete treatment…

Artificial Intelligence · Computer Science 2025-01-22 Cheuk Hang Leung , Yiyan Huang , Yijun Li , Qi Wu

Recent research has shown the existence of significant redundancy in large Transformer models. One can prune the redundant parameters without significantly sacrificing the generalization performance. However, we question whether the…

Computation and Language · Computer Science 2022-02-15 Chen Liang , Haoming Jiang , Simiao Zuo , Pengcheng He , Xiaodong Liu , Jianfeng Gao , Weizhu Chen , Tuo Zhao

Explicit finite-sample statistical guarantees on model performance are an important ingredient in responsible machine learning. Previous work has focused mainly on bounding either the expected loss of a predictor or the probability that an…

Machine Learning · Computer Science 2024-03-07 Zhun Deng , Thomas P. Zollo , Jake C. Snell , Toniann Pitassi , Richard Zemel

This paper provides a formal econometric framework behind the newly developed difference-in-discontinuities design (DiDC). Despite its increasing use in applied research, there are currently limited studies of its properties. We formalize…

Econometrics · Economics 2026-01-28 Pedro Picchetti , Cristine C. X. Pinto , Stephanie T. Shinoki

In the statistical analysis of objects, samples and populations with quantitative variables, in many occasions we are interested in knowing the proportions that exist between the different variables from a same object; if these proportions…

Data Analysis, Statistics and Probability · Physics 2007-05-23 Enrique Ordaz Romay

For many machine learning problems, data is abundant and it may be prohibitive to make multiple passes through the full training set. In this context, we investigate strategies for dynamically increasing the effective sample size, when…

Machine Learning · Computer Science 2016-10-10 Hadi Daneshmand , Aurelien Lucchi , Thomas Hofmann

In health cohort studies, repeated measures of markers are often used to describe the natural history of a disease. Joint models allow to study their evolution by taking into account the possible informative dropout usually due to clinical…

Modal regression has emerged as a flexible alternative to classical regression models when the conditional mean or median are unable to adequately capture the underlying relation between a response and a predictor variable. This approach is…

Methodology · Statistics 2025-04-08 Ana Pérez-González , Tomás R. Cotos-Yáñez , Rosa M. Crujeiras

Video anomaly detection is a challenging task due to the lack in approaches for representing samples. The visual representations of most existing approaches are limited by short-term sequences of observations which cannot provide enough…

Computer Vision and Pattern Recognition · Computer Science 2023-09-08 Yalong Jiang , Changkang Li

A fundamental issue in causal inference for Big Observational Data is confounding due to covariate imbalances between treatment groups. This can be addressed by designing the data prior to analysis. Existing design methods, developed for…

Methodology · Statistics 2022-03-17 Yumin Zhang , Arman Sabbaghi

A basic principle in the design of observational studies is to approximate the randomized experiment that would have been conducted under controlled circumstances. Now, linear regression models are commonly used to analyze observational…

Methodology · Statistics 2022-07-08 Ambarish Chattopadhyay , Jose R. Zubizarreta

Matching is a commonly used causal inference study design in observational studies. Through matching on measured confounders between different treatment groups, valid randomization inferences can be conducted under the no unmeasured…

Methodology · Statistics 2024-09-20 Jeffrey Zhang , Siyu Heng

We consider the sequential experimental design problem in the predict-then-optimize paradigm. In this paradigm, the outputs of the prediction model are used as coefficient vectors in a downstream linear optimization problem. Traditional…

Machine Learning · Statistics 2026-02-06 Beichen Wan , Mo Liu , Paul Grigas , Zuo-Jun Max Shen

We propose a framework for building patient-specific treatment recommendation models, building on the large recent literature on learning patient-level causal models and inspired by the target trial paradigm of Hernan and Robins. We focus…

Machine Learning · Statistics 2025-07-17 Rom Gutman , Shimon Sheiba , Omer Noy Klein , Naama Dekel Bird , Amit Gruber , Doron Aronson , Oren Caspi , Uri Shalit

Recent work introduced loss functions which measure the error of a prediction based on multiple simultaneous observations or outcomes. In this paper, we explore the theoretical and practical questions that arise when using such…

Machine Learning · Computer Science 2018-02-28 Rafael Frongillo , Nishant A. Mehta , Tom Morgan , Bo Waggoner

Session-based Recommendation (SR) aims to predict the next item for recommendation based on previously recorded sessions of user interaction. The majority of existing approaches to SR focus on modeling the transition patterns of items. In…

Information Retrieval · Computer Science 2022-04-06 Jiahao Yuan , Wendi Ji , Dell Zhang , Jinwei Pan , Xiaoling Wang

The Poisson distribution is the default choice of likelihood for probabilistic models of count data. However, due to the equidispersion contraint of the Poisson, such models may have predictive uncertainty that is artificially inflated.…

Methodology · Statistics 2025-07-15 Jimmy Lederman , Aaron Schein

This report continues our investigation of effects a simulation design may have on the conclusions on performance of statistical methods. In the context of meta-analysis of log-odds-ratios, we consider five generation mechanisms for control…

Methodology · Statistics 2020-07-13 Elena Kulinskaya , David C. Hoaglin , Ilyas Bakbergenuly