English
Related papers

Related papers: On Efficient Design of Pilot Experiment for Genera…

200 papers

Generalized linear models (GLMs) arise in high-dimensional machine learning, statistics, communications and signal processing. In this paper we analyze GLMs when the data matrix is random, as relevant in problems such as compressed sensing,…

Information Theory · Computer Science 2019-04-01 Jean Barbier , Florent Krzakala , Nicolas Macris , Léo Miolane , Lenka Zdeborová

The selection of optimal designs for generalized linear mixed models is complicated by the fact that the Fisher information matrix, on which most optimality criteria depend, is computationally expensive to evaluate. Our focus is on the…

Methodology · Statistics 2015-09-22 Timothy W. Waite , David C. Woods

Optimal experiment design for parameter estimation is a research topic that has been in the interest of various studies. A key problem in optimal input design is that the optimal input depends on some unknown system parameters that are to…

Systems and Control · Computer Science 2019-04-17 Lirong Huang , Håkan Hjalmarsson , László Gerencsér

Generalized Linear Models (GLMs) and Single Index Models (SIMs) provide powerful generalizations of linear regression, where the target variable is assumed to be a (possibly unknown) 1-dimensional function of a linear predictor. In general,…

Artificial Intelligence · Computer Science 2011-04-14 Sham Kakade , Adam Tauman Kalai , Varun Kanade , Ohad Shamir

Generalized linear models (GLMs) have been used widely for modelling the mean response both for discrete and continuous random variables with an emphasis on categorical response. Recently Yang, Mandal and Majumdar (2013) considered full…

Statistics Theory · Mathematics 2013-05-06 Jie Yang , Abhyuday Mandal

Empirical research plays a fundamental role in the machine learning domain. At the heart of impactful empirical research lies the development of clear research hypotheses, which then shape the design of experiments. The execution of…

Machine Learning · Computer Science 2024-05-29 Daniel Vranješ , Oliver Niggemann

Pre-experiment stratification, or blocking, is a well-established technique for designing more efficient experiments and increasing the precision of the experimental estimates. However, when researchers have access to many covariates at the…

Econometrics · Economics 2025-10-01 George Gui , Seungwoo Kim

A computer model can be used for predicting an output only after specifying the values of some unknown physical constants known as calibration parameters. The unknown calibration parameters can be estimated from real data by conducting…

Methodology · Statistics 2021-06-18 Arvind Krishna , V. Roshan Joseph , Shan Ba , William A. Brenneman , William R. Myers

Randomized response (RR) designs are used to collect response data about sensitive behaviors (e.g., criminal behavior, sexual desires). The modeling of RR data is more complex, since it requires a description of the RR process. For the…

Methodology · Statistics 2021-06-21 Jean-Paul Fox , Konrad Klotzke , Duco Veen

Selective inference aims at providing valid inference after a data-driven selection of models or hypotheses. It is essential to avoid overconfident results and replicability issues. While significant advances have been made in this area for…

Methodology · Statistics 2025-03-14 Matteo D'Alessandro , Magne Thoresen

We consider the problem of robust inference under the generalized linear model (GLM) with stochastic covariates. We derive the properties of the minimum density power divergence estimator of the parameters in GLM with random design and use…

Methodology · Statistics 2020-04-06 Ayanendranath Basu , Abhik Ghosh , Abhijit Mandal , Nirian Martin , Leandro Pardo

Large language models (LLMs) have recently been proposed as general-purpose agents for experimental design, with claims that they can perform in-context experimental design. We evaluate this hypothesis using both open- and closed-source…

Machine Learning · Computer Science 2025-09-29 Rushil Gupta , Jason Hartford , Bang Liu

How should researchers analyze randomized experiments in which the main outcome is latent and measured in multiple ways but each measure contains some degree of error? We first identify a critical study-specific noncomparability problem in…

Econometrics · Economics 2026-01-13 Jiawei Fu , Donald P. Green

This paper focuses on modern efficient training and inference technologies on foundation models and illustrates them from two perspectives: model and system design. Model and System Design optimize LLM training and inference from different…

Distributed, Parallel, and Cluster Computing · Computer Science 2025-04-15 Dong Liu , Yanxuan Yu , Yite Wang , Jing Wu , Zhongwei Wan , Sina Alinejad , Benjamin Lengerich , Ying Nian Wu

Despite rapid progress in large language models (LLMs), the statistical structure of their weights, activations, and gradients-and its implications for initialization, training dynamics, and efficiency-remains largely unexplored. We…

Machine Learning · Computer Science 2026-02-24 Jun Wu , Patrick Huang , Jiangtao Wen , Yuxing Han

In this paper, we address the problem of designing an experimental plan with both discrete and continuous factors under fairly general parametric statistical models. We propose a new algorithm, named ForLion, to search for locally optimal…

Computation · Statistics 2024-05-24 Yifei Huang , Keren Li , Abhyuday Mandal , Jie Yang

Generalized Linear Models (GLMs) are an increasingly popular framework for modeling neural spike trains. They have been linked to the theory of stochastic point processes and researchers have used this relation to assess goodness-of-fit…

Neurons and Cognition · Quantitative Biology 2010-11-19 Felipe Gerhard , Wulfram Gerstner

Model checking is essential to evaluate the adequacy of statistical models and the validity of inferences drawn from them. Particularly, hierarchical models such as latent Gaussian models (LGMs) pose unique challenges as it is difficult to…

Methodology · Statistics 2023-07-25 Rafael Cabral , David Bolin , Håvard Rue

Large language models (LLMs) are increasingly used to simulate human behavior, but common practices to use LLM-generated data are inefficient. Treating an LLM's output ("model choice") as a single data point underutilizes the information…

Artificial Intelligence · Computer Science 2025-12-30 Hongshen Sun , Juanjuan Zhang

Finite mixture distributions arise in sampling a heterogeneous population. Data drawn from such a population will exhibit extra variability relative to any single subpopulation. Statistical models based on finite mixtures can assist in the…

Methodology · Statistics 2024-01-19 Andrew M. Raim , Nagaraj K. Neerchal , Jorge G. Morel