English
Related papers

Related papers: On Spatial Lag Models estimated using crowdsourcin…

200 papers

Traffic congestion research is on the rise, thanks to urbanization, economic growth, and industrialization. Developed countries invest a lot of research money in collecting traffic data using Radio Frequency Identification (RFID), loop…

Class imbalance and distributional differences in large datasets present significant challenges for classification tasks machine learning, often leading to biased models and poor predictive performance for minority classes. This work…

Machine Learning · Statistics 2024-12-20 Alex Mak , Shubham Sahoo , Shivani Pandey , Yidan Yue , Linglong Kong

We develop a new sampling method to estimate eigenvector centrality on incomplete networks. Our goal is to estimate this global centrality measure having at disposal a limited amount of data. This is the case in many real-world scenarios…

Social and Information Networks · Computer Science 2020-10-29 Nicolò Ruggeri , Caterina De Bacco

Customer scoring models are the core of scalable direct marketing. Uplift models provide an estimate of the incremental benefit from a treatment that is used for operational decision-making. Training and monitoring of uplift models require…

Machine Learning · Computer Science 2019-10-02 Johannes Haupt , Daniel Jacob , Robin M. Gubela , Stefan Lessmann

In this paper we describe two bootstrap methods for massive data sets. Naive applications of common resampling methodology are often impractical for massive data sets due to computational burden and due to complex patterns of inhomogeneity.…

Applications · Statistics 2013-01-14 S. N. Lahiri , C. Spiegelman , J. Appiah , L. Rilett

High resolution geospatial data are challenging because standard geostatistical models based on Gaussian processes are known to not scale to large data sizes. While progress has been made towards methods that can be computed more…

Methodology · Statistics 2020-12-03 Michele Peruzzi , David B. Dunson

Predictive variability due to data ambiguities has typically been addressed via construction of dedicated models with built-in probabilistic capabilities that are trained to predict uncertainty estimates as variables of interest. These…

Machine Learning · Computer Science 2023-08-04 Katarína Tóthová , Ľubor Ladický , Daniel Thul , Marc Pollefeys , Ender Konukoglu

In many randomized trials, outcomes such as essays or open-ended responses must be manually scored as a preliminary step to impact analysis, a process that is costly and limiting. Model-assisted estimation offers a way to combine surrogate…

Methodology · Statistics 2026-02-16 Reagan Mozer , Nicole E. Pashley , Luke Miratrix

We investigate spatial confounding in the presence of multivariate disease dependence. In the "analysis model perspective" of spatial confounding, adding a spatially dependent random effect can lead to significant variance inflation of the…

Methodology · Statistics 2026-02-13 Kyle Lin Wu , Sudipto Banerjee

This paper investigates the use of stratified sampling as a variance reduction technique for approximating integrals over large dimensional spaces. The accuracy of this method critically depends on the choice of the space partition, the…

Probability · Mathematics 2009-09-15 Pierre Etoré , Gersende Fort , Benjamin Jourdain , Eric Moulines

Large-scale networks are commonly encountered in practice (e.g., Facebook and Twitter) by researchers. In order to study the network interaction between different nodes of large-scale networks, the spatial autoregressive (SAR) model has…

Computation · Statistics 2023-06-12 Xuetong Li , Feifei Wang , Wei Lan , Hansheng Wang

Data privacy has increasingly become a daunting challenge because it limits data availability, which is essential in estimating statistical models such as generalized linear mixed models. Access to personal data often involves considerable…

Methodology · Statistics 2026-05-05 Marie Analiz April Limpoco , Christel Faes , Niel Hens

Spatial range joins have many applications, including geographic information systems, location-based social networking services, neuroscience, and visualization. However, joins incur not only expensive computational costs but also too large…

Databases · Computer Science 2025-08-22 Daichi Amagata

We investigate the use of a stratified sampling approach for LIME Image, a popular model-agnostic explainable AI method for computer vision tasks, in order to reduce the artifacts generated by typical Monte Carlo sampling. Such artifacts…

Artificial Intelligence · Computer Science 2024-03-27 Muhammad Rashid , Elvio G. Amparore , Enrico Ferrari , Damiano Verda

This paper proposes a new Sequential Monte Carlo algorithm to perform online estimation in the context of state space models when either the transition density of the latent state or the conditional likelihood of an observation given a…

Applications · Statistics 2021-05-10 Alice Martin , Marie-Pierre Etienne , Pierre Gloaguen , Sylvain Le Corff , Jimmy Olsson

Statistical models with constrained probability distributions are abundant in machine learning. Some examples include regression models with norm constraints (e.g., Lasso), probit, many copula models, and latent Dirichlet allocation (LDA).…

Computation · Statistics 2015-06-22 Shiwei Lan , Babak Shahbaba

Spatial generalized linear mixed models (SGLMMs) are popular and flexible models for non-Gaussian spatial data. They are useful for spatial interpolations as well as for fitting regression models that account for spatial dependence, and are…

Methodology · Statistics 2021-10-26 Yawen Guan , Murali Haran

The application of state-of-the-art spatial econometric models requires that the information about the spatial coordinates of statistical units is completely accurate, which is usually the case in the context of areal data. With…

Methodology · Statistics 2019-09-06 Giuseppe Arbia , Maria Michela Dickson , Giuseppe Espa , Diego Giuliani , Flavio Santi

Data plays a fundamental role in the training of Large Language Models (LLMs). While attention has been paid to the collection and composition of datasets, determining the data sampling strategy in training remains an open question. Most…

Computation and Language · Computer Science 2024-06-04 Yunfan Shao , Linyang Li , Zhaoye Fei , Hang Yan , Dahua Lin , Xipeng Qiu

In this paper we adapt online estimation strategies to perform model-based clustering on large networks. Our work focuses on two algorithms, the first based on the SAEM algorithm, and the second on variational methods. These two strategies…

Applications · Statistics 2010-11-11 Hugo Zanghi , Franck Picard , Vincent Miele , Christophe Ambroise