中文
相关论文

相关论文: Computationally efficient univariate filtering for…

200 篇论文

Large-scale retrieval systems are often implemented as a cascading sequence of phases -- a first filtering step, in which a large set of candidate documents are extracted using a simple technique such as Boolean matching and/or static…

信息检索 · 计算机科学 2015-06-03 Charles L. A. Clarke , J. Shane Culpepper , Alistair Moffat

The paper considers variable selection in linear regression models where the number of covariates is possibly much larger than the number of observations. High dimensionality of the data brings in many complications, such as (possibly…

统计方法学 · 统计学 2016-11-29 Haeran Cho , Piotr Fryzlewicz

Small study effects occur when smaller studies show different, often larger, treatment effects than large ones, which may threaten the validity of systematic reviews and meta-analyses. The most well-known reasons for small study effects…

统计方法学 · 统计学 2018-05-28 Chuan Hong , Georgia Salanti , Sally Morton , Richard Riley , Haitao Chu , Stephen E. Kimmel , Yong Chen

We propose a generic Markov Chain Monte Carlo (MCMC) algorithm to speed up computations for datasets with many observations. A key feature of our approach is the use of the highly efficient difference estimator from the survey sampling…

统计方法学 · 统计学 2017-08-03 Matias Quiroz , Mattias Villani , Robert Kohn

A popular approach for testing if two univariate random variables are statistically independent consists of partitioning the sample space into bins, and evaluating a test statistic on the binned data. The partition size matters, and the…

统计方法学 · 统计学 2016-04-28 Ruth Heller , Yair Heller , Shachar Kaufman , Barak Brill , Malka Gorfine

We describe how to calculate standard errors for A/B tests that include clustered data, ratio metrics, and/or covariate adjustment. We may do this for power analysis/sample size calculations prior to running an experiment using historical…

统计方法学 · 统计学 2024-06-12 Tim Hesterberg , Ben Knight

We present a scalable framework for computing polygenic risk scores (PRS) in high-dimensional genomic settings using the recently introduced Univariate-Guided Sparse Regression (uniLasso). UniLasso is a two-stage penalized regression…

Subsampling algorithms for various parametric regression models with massive data have been extensively investigated in recent years. However, all existing studies on subsampling heavily rely on clean massive data. In practical…

统计理论 · 数学 2025-06-11 Jiangshan Ju , Mingqiu Wang , Shengli Zhao

Within ab initio Quantum Monte Carlo simulations, the leading numerical cost for large systems is the computation of the values of the Slater determinants in the trial wavefunction. Each Monte Carlo step requires finding the determinant of…

计算物理 · 物理学 2017-11-22 T. McDaniel , E. F. D'Azevedo , Y. W. Li , K. Wong , P. R. C. Kent

This paper introduces an open-ended sequential algorithm for computing the p-value of a test using Monte Carlo simulation. It guarantees that the resampling risk, the probability of a different decision than the one based on the theoretical…

统计理论 · 数学 2013-07-30 Axel Gandy

Random sampling has become a critical tool in solving massive matrix problems. For linear regression, a small, manageable set of data rows can be randomly selected to approximate a tall, skinny data matrix, improving processing time…

数据结构与算法 · 计算机科学 2014-08-22 Michael B. Cohen , Yin Tat Lee , Cameron Musco , Christopher Musco , Richard Peng , Aaron Sidford

In variable selection, most existing screening methods focus on marginal effects and ignore dependence between covariates. To improve the performance of selection, we incorporate pairwise effects in covariates for screening and…

统计方法学 · 统计学 2019-02-12 Siliang Gong , Kai Zhang , Yufeng Liu

Calibration tests based on the probability integral transform (PIT) are routinely used to assess the quality of univariate distributional forecasts. However, PIT-based calibration tests for multivariate distributional forecasts face various…

计量经济学 · 经济学 2023-12-13 Malte Knüppel , Fabian Krüger , Marc-Oliver Pohle

The use of the Monte Carlo technique in a reliable and inexpensive way without the need for a standard radioactive source in determining the detector efficiency is becoming widespread every passing day. It is important to model the detector…

仪器与探测器 · 物理学 2023-01-18 Esra Uyar , Zeynep Aybüke Günekbay

Performance measurement is an essential task once a statistical model is created. The Area Under the receiving operating characteristics Curve (AUC) is the most popular measure for evaluating the quality of a binary classifier. In this…

统计计算 · 统计学 2021-05-24 Robin Van Oirbeek , Jolien Ponnet , Tim Verdonck

The computational cost for inference and prediction of statistical models based on Gaussian processes with Mat\'ern covariance functions scales cubicly with the number of observations, limiting their applicability to large data sets. The…

统计理论 · 数学 2025-03-04 David Bolin , Vaibhav Mehandiratta , Alexandre B. Simas

Gaussian process regression is a frequently used statistical method for flexible yet fully probabilistic non-linear regression modeling. A common obstacle is its computational complexity which scales poorly with the number of observations.…

统计方法学 · 统计学 2026-03-10 Adam Gorm Hoffmann , Claus Thorn Ekstrøm , Andreas Kryger Jensen

Much work has been dedicated to estimating and optimizing workloads in high-performance computing (HPC) and deep learning. However, researchers have typically relied on few metrics to assess the efficiency of those techniques. Most notably,…

机器学习 · 计算机科学 2023-10-17 Hugo Waltsburger , Erwan Libessart , Chengfang Ren , Anthony Kolar , Regis Guinvarc'h

In comparative studies, such as in causal inference and clinical trials, balancing important covariates is often one of the most important concerns for both efficient and credible comparison. However, chance imbalance still exists in many…

统计方法学 · 统计学 2018-07-30 Yichen Qin , Yang Li , Wei Ma , Feifang Hu

In many industries, predicting metric outcomes of large systems is a fundamental problem, driven largely by traditional tabular regression. However, such methods struggle on complex systems data in the wild such as configuration files or…