Related papers: Heterogeneous Elasticities, Aggregation, and Retra…
The output of predictive models is routinely recalibrated by reconciling low-level predictions with known derived quantities defined at higher levels of aggregation. For example, models predicting turnout probabilities at the individual…
Meta-learning aims to leverage information across related tasks to improve prediction on unlabeled data for new tasks when only a small number of labeled observations are available ("few-shot" learning). Increased task diversity is often…
We consider a static linear panel model with both correlated and uncorrelated random coefficients, where the former can depend arbitrarily on observable regressors while the latter are independent of them. We provide sufficient conditions…
Ordinary least-squares (OLS) estimators for a linear model are very sensitive to unusual values in the design space or outliers among y values. Even one single atypical value may have a large effect on the parameter estimates. This article…
Multivariate binary distributions can be decomposed into products of univariate conditional distributions. Recently popular approaches have modeled these conditionals through neural networks with sophisticated weight-sharing structures. It…
Elastic constants and mechanical properties play a pivotal role across multiple disciplines and engineering applications. We introduced the optimized high-efficient strain-matrix set (OHESS) that determines the second-order elastic…
Selection bias arises when the probability that an observation enters a dataset depends on variables related to the quantities of interest, leading to systematic distortions in estimation and uncertainty quantification. For example, in…
We study the elasto-plastic behaviour of materials made of individual (discrete) objects, such as a liquid foam made of bubbles. The evolution of positions and mutual arrangements of individual objects is taken into account through…
Which parts of a dataset will a given model find difficult? Recent work has shown that SGD-trained models have a bias towards simplicity, leading them to prioritize learning a majority class, or to rely upon harmful spurious correlations.…
Feature selection is important in data representation and intelligent diagnosis. Elastic net is one of the most widely used feature selectors. However, the features selected are dependant on the training data, and their weights dedicated…
Logistic models are studied as a tool to convert output from numerical weather forecasting systems (deterministic and ensemble) into probability forecasts for binary events. A logistic model obtains by putting the logarithmic odds ratio…
A data set sampled from a certain population is biased if the subgroups of the population are sampled at proportions that are significantly different from their underlying proportions. Training machine learning models on biased data sets…
The elastic behavior of materials operating in the linear regime is constrained, by definition, to operations that are linear in the imposed deformation. Though the nonlinear regime holds promise for new functionality, the design in this…
The coefficient of normal restitution in an oblique impact is theoretically studied. Using a two-dimensional lattice models for an elastic disk and an elastic wall, we demonstrate that the coefficient of normal restitution can exceed one…
Logistic regression models are widely used in the social and behavioral sciences and in high-stakes domains, due to their simplicity and interpretability properties. At the same time, such domains are permeated by distribution shifts, where…
Multi-step forecasting is often described through a simple rule of thumb: recursive strategies are said to have high bias and low variance, while direct strategies are said to have low bias and high variance. We revisit this belief by…
Mechanical deformation of amorphous solids can be described as consisting of an elastic part in which the stress increases linearly with strain, up to a yield point at which the solid either fractures or starts deforming plastically. It is…
Consider a researcher estimating the parameters of a regression function based on data for all 50 states in the United States or on data for all visits to a website. What is the interpretation of the estimated parameters and the standard…
This paper looks at effects, due to the boundary, on inference in logistic regression. It shows that first -- and, indeed, higher -- order asymptotic results are not uniform across the model. Near the boundary, effects such as high…
The log odds ratio is a well-established metric for evaluating the association between binary outcome and exposure variables. Despite its widespread use, there has been limited discussion on how to summarize the log odds ratio as a function…