Related papers: Sufficient Conditions for the Shrinking Wellness L…
Through the lense of multilevel model (MLM) specification and regularization, this is a connect-the-dots introductory summary of Small Area Estimation, e.g. small group prediction informed by a complex sampling design. While a comprehensive…
Many set selection and ranking algorithms have recently been enhanced with diversity constraints that aim to explicitly increase representation of historically disadvantaged populations, or to improve the overall representativeness of the…
We introduce a new sufficient statistic for the population parameter vector by allowing for the sampling design to first be selected at random amongst a set of candidate sampling designs. In contrast to the traditional approach in survey…
For every $\alpha \leq \beta$ in a left neighborhood $[\alpha_0,1]$ of 1, a group $G(\alpha,\beta)$ is constructed, the growth function of which satisfies $\limsup \frac{\log \log b_{G(\alpha,\beta)}(r)}{\log r}=\alpha$ and $\liminf…
Gnutzmann and Zyczkowski have proposed the Renyi-Wehrl entropy as a generalization of the Wehrl entropy, and conjectured that its minimum is obtained for coherent states. We prove this conjecture for the Renyi index q=2,3,... in the cases…
Reinforcement learning with verifiable rewards (RLVR) has become a standard recipe for improving large language models (LLMs) on reasoning tasks, with Group Relative Policy Optimization (GRPO) widely used in practice. Yet GRPO wastes…
We provide (high probability) bounds on the condition number of random feature matrices. In particular, we show that if the complexity ratio $\frac{N}{m}$ where $N$ is the number of neurons and $m$ is the number of data samples scales like…
We consider panel data models with group structure. We study the asymptotic behavior of least-squares estimators and information criterion for the number of groups, allowing for the presence of small groups that have an asymptotically…
Internal measures that are used to assess the quality of a clustering usually take into account intra-group and/or inter-group criteria. There are many papers in the literature that propose algorithms with provable approximation guarantees…
Large Language Models (LLMs) are advancing rapidly, yet the benchmarks used to measure this progress are becoming increasingly unreliable. Score inflation and selective reporting have eroded the authority of standard benchmarks, leaving the…
Reinforcement Learning (RL) has emerged as a highly effective technique for addressing various scientific and applied problems. Despite its success, certain complex tasks remain challenging to be addressed solely with a single model and…
We study sparse hypergraphs which satisfy a mild pseudorandomness condition known as $L_p$ regularity. We prove appropriate regularity and counting lemmas, and we extend the relative removal lemma of Tao in this setting. This answers a…
In group sequential analysis, data is collected and analyzed in batches until pre-defined stopping criteria are met. Inference in the parametric setup typically relies on the limiting asymptotic multivariate normality of the repeatedly…
Group testing is a well-known search problem that consists in detecting of $s$ defective members of a set of $t$ samples by carrying out tests on properly chosen subsets of samples. In classical group testing the goal is to find all…
Millions of individuals' well-being are challenged by the harms of substance use. Harm reduction as a public health strategy is designed to improve their health outcomes and reduce safety risks. Some large language models (LLMs) have…
Despite widespread adoption in practice, guarantees for the LASSO and Group LASSO are strikingly lacking in settings beyond statistical problems, and these algorithms are usually considered to be a heuristic in the context of sparse convex…
Spatial clustering has important implications in various fields. In particular, disease clustering is of major public concern in epidemiology. In this article, we propose the use of two distance-based segregation indices to test the…
In fair machine learning, one source of performance disparities between groups is over-fitting to groups with relatively few training samples. We derive group-specific bounds on the generalization error of welfare-centric fair machine…
Many special classes of simplicial sets, such as the nerves of categories or groupoids, the 2-Segal sets of Dyckerhoff and Kapranov, and the (discrete) decomposition spaces of G\'{a}lvez, Kock, and Tonks, are characterized by the property…
We consider a size-structured aggregation and growth model of phytoplankton community proposed by Ackleh and Fitzpatrick [2]. The model accounts for basic biological phenomena in phytoplankton community such as growth, gravitational…