English
Related papers

Related papers: The Effects of Data Size and Frequency Range on Di…

200 papers

Marginal structural models fit via inverse probability of treatment weighting are commonly used to control for confounding when estimating causal effects from observational data. When planning a study that will be analyzed with marginal…

Applications · Statistics 2020-03-16 Bonnie E. Shook-Sa , Michael G. Hudgens

The purpose of this project was to collect and analyse data about the comparability and real-life applicability of published results focusing on Microsoft Windows malware, more specifically the impact of dataset size and testing dataset…

Cryptography and Security · Computer Science 2022-06-14 David Illes

Current explanation techniques towards a transparent Convolutional Neural Network (CNN) mainly focuses on building connections between the human-understandable input features with models' prediction, overlooking an alternative…

Machine Learning · Computer Science 2020-05-08 Zifan Wang , Yilin Yang , Ankit Shrivastava , Varun Rawal , Zihao Ding

This paper examines the efficiency of spatial and frequency dimensions in serving multiple users in the downlink of a small cell wireless network with randomly deployed access points. For this purpose, the stochastic geometry framework is…

Information Theory · Computer Science 2014-02-18 Stelios Stefanatos , Angeliki Alexiou

Recent studies have shown that diffusion language models achieve remarkable data efficiency under limited-data constraints, yet the underlying mechanisms remain unclear. In this work, we perform extensive ablation experiments to disentangle…

Computation and Language · Computer Science 2025-10-07 Zitian Gao , Haoming Luo , Lynx Chen , Jason Klein Liu , Ran Tao , Joey Zhou , Bryan Dai

We develop randomization-based tests for heterogeneous treatment effects in the presence of network interference. Leveraging the exposure mapping framework, we study a broad class of null hypotheses that represent various forms of constant…

Econometrics · Economics 2025-06-25 Julius Owusu

We quantify the finite size effects in a stochastic network made up of rate neurons, for several kinds of recurrent connectivity matrices. This analysis is performed by means of a perturbative expansion of the neural equations, where the…

Dynamical Systems · Mathematics 2013-07-09 D. Fasoli , O. Faugeras

A basic issue in both teaching of and practice of statistics is the interplay between modelling assumptions and inference performance. The general message conveyed is that stronger assumptions lead to better statistical performance of the…

Statistics Theory · Mathematics 2026-03-20 Morten Byholt , Nils Lid Hjort

The problem tackled in this paper is the determination of sample size for a given level and power in the context of a simple linear regression model. At a technical level, the simple linear regression model is a five-parameter model. It is…

Methodology · Statistics 2019-07-25 Tianyuan Guan , M. Khorshed Alam , M. Bhaskara Rao

This paper investigates the impact of shallow versus deep relevance judgments on the performance of BERT-based reranking models in neural Information Retrieval. Shallow-judged datasets, characterized by numerous queries each with few…

Information Retrieval · Computer Science 2025-07-01 Gabriel Iturra-Bocaz , Danny Vo , Petra Galuscakova

It is widely acknowledged that trained convolutional neural networks (CNNs) have different levels of sensitivity to signals of different frequency. In particular, a number of empirical studies have documented CNNs sensitivity to…

Machine Learning · Computer Science 2023-09-27 Charles Godfrey , Elise Bishoff , Myles Mckay , Davis Brown , Grayson Jorgenson , Henry Kvinge , Eleanor Byler

`Distribution regression' refers to the situation where a response Y depends on a covariate P where P is a probability distribution. The model is Y=f(P) + mu where f is an unknown regression function and mu is a random error. Typically, we…

Machine Learning · Statistics 2013-02-04 Barnabas Poczos , Alessandro Rinaldo , Aarti Singh , Larry Wasserman

A systematic, comparative investigation into the effects of low-quality data reveals a stark spectrum of robustness across modern probabilistic models. We find that autoregressive language models, from token prediction to…

Artificial Intelligence · Computer Science 2025-12-16 Liu Peng , Yaochu Jin

Understanding geometric properties of natural language processing models' latent spaces allows the manipulation of these properties for improved performance on downstream tasks. One such property is the amount of data spread in a model's…

Machine Learning · Computer Science 2023-08-02 Anna C. Marbut , Katy McKinney-Bock , Travis J. Wheeler

Compressing self-supervised models has become increasingly necessary, as self-supervised models become larger. While previous approaches have primarily focused on compressing the model size, shortening sequences is also effective in…

Computation and Language · Computer Science 2022-10-26 Yen Meng , Hsuan-Jui Chen , Jiatong Shi , Shinji Watanabe , Paola Garcia , Hung-yi Lee , Hao Tang

With the violation of the assumption of homoskedasticity, least squares estimators of the variance become inefficient and statistical inference conducted with invalid standard errors leads to misleading rejection rates. Despite a vast…

Econometrics · Economics 2024-01-01 Annalivia Polselli

The general relationship between an arbitrary frequency distribution and the expectation value of the frequency distributions of its samples is discussed. A wide set of measurable quantities ("invariant moments") whose expectation value…

Data Analysis, Statistics and Probability · Physics 2015-06-15 Paolo Rossi

Distribution regression refers to the supervised learning problem where labels are only available for groups of inputs instead of individual inputs. In this paper, we develop a rigorous mathematical framework for distribution regression…

Machine Learning · Computer Science 2021-09-30 Maud Lemercier , Cristopher Salvi , Theodoros Damoulas , Edwin V. Bonilla , Terry Lyons

Two semimetrics on probability distributions are proposed, given as the sum of differences of expectations of analytic functions evaluated at spatial or frequency locations (i.e, features). The features are chosen so as to maximize the…

Machine Learning · Statistics 2016-10-31 Wittawat Jitkrittum , Zoltan Szabo , Kacper Chwialkowski , Arthur Gretton

The dependency of the generalization error of neural networks on model and dataset size is of critical importance both in practice and for understanding the theory of neural networks. Nevertheless, the functional form of this dependency…

Machine Learning · Computer Science 2019-12-23 Jonathan S. Rosenfeld , Amir Rosenfeld , Yonatan Belinkov , Nir Shavit