中文
相关论文

相关论文: `pandemonium`: High Dimensional Analysis in Linked…

200 篇论文

Technology and collaboration enable dramatic increases in the size of psychological and psychiatric data collections, but finding structure in these large data sets with many collected variables is challenging. Decision tree ensembles like…

机器学习 · 统计学 2017-02-15 Patrick J. Miller , Gitta H. Lubke , Daniel B. McArtor , C. S. Bergeman

Motivation: Model selection is a ubiquitous challenge in statistics. For penalized models, model selection typically entails tuning hyperparameters to maximize a measure of fit or minimize out-of-sample prediction error. However, these…

统计方法学 · 统计学 2025-05-29 Priyam Das , Sarah Robinson , Christine B. Peterson

Analyzing time-series cross-sectional (also known as longitudinal or panel) data is an important process across a number of fields, including the social sciences, economics, finance, and medicine. PanelMatch is an R package that implements…

统计方法学 · 统计学 2025-08-19 Adam Rauh , In Song Kim , Kosuke Imai

We provide a pipeline for calculating, managing and visualising correlations and other pairwise association scores for numerical and categorical data. We present a uniform interface for calculating a plethora of pairwise scores and propose…

统计计算 · 统计学 2025-11-17 Amit Chinwan , Catherine B. Hurley

Deep clustering incorporates embedding into clustering to find a lower-dimensional space appropriate for clustering. In this paper, we propose a novel deep clustering framework with self-supervision using pairwise similarities (DCSS). The…

机器学习 · 计算机科学 2024-05-07 Mohammadreza Sadeghi , Narges Armanfard

In multiple correspondence analysis, both individuals (observations) and categories can be represented in a biplot that jointly depicts the relationships across categories or individuals, as well as the associations between them. Additional…

统计方法学 · 统计学 2019-01-10 Mariko Takagishi , Michel van de Velden

When some 'entities' are related by the 'features' they share they are amenable to a bipartite network representation. Plant-pollinator ecological communities, co-authorship of scientific papers, customers and purchases, or answers in a…

社会与信息网络 · 计算机科学 2020-10-14 Ignacio Tamarit , María Pereda , José A. Cuesta

In this work we present a statistical approach to distinguish and interpret the complex relationship between several predictors and a response variable at the small area level, in the presence of i) high correlation between the predictors…

应用统计 · 统计学 2016-02-24 Silvia Liverani , Aurore Lavigne , Marta Blangiardo

The application of large language models (LLMs) in clinical decision support faces significant challenges of "tunnel vision" and diagnostic hallucinations present in their processing unstructured electronic health records (EHRs). To address…

人工智能 · 计算机科学 2026-04-28 Zhiqi Lv , Duofan Tu , Jun Li , Mingyue Zhao , Heqin Zhu , Wenliang Li , Shaohua Kevin Zhou

Evaluating forecasts is essential to understand and improve forecasting and make forecasts useful to decision makers. A variety of R packages provide a broad variety of scoring rules, visualisations and diagnostic tools. One particular…

统计方法学 · 统计学 2024-11-04 Nikos I. Bosse , Hugo Gruson , Anne Cori , Edwin van Leeuwen , Sebastian Funk , Sam Abbott

Graphical models provide powerful tools to uncover complicated patterns in multivariate data and are commonly used in Bayesian statistics and machine learning. In this paper, we introduce the R package BDgraph which performs Bayesian…

机器学习 · 统计学 2019-05-14 Reza Mohammadi , Ernst C. Wit

Learning graphical models from data is an important problem with wide applications, ranging from genomics to the social sciences. Nowadays datasets often have upwards of thousands---sometimes tens or hundreds of thousands---of variables and…

机器学习 · 统计学 2019-11-26 Bryon Aragam , Jiaying Gu , Qing Zhou

Pandemic measures such as social distancing and contact tracing can be enhanced by rapidly integrating dynamic location data and demographic data. Projecting billions of longitude and latitude locations onto hundreds of thousands of highly…

Clustering algorithms are one of the main analytical methods to detect patterns in unlabeled data. Existing clustering methods typically treat samples in a dataset as points in a metric space and compute distances to group together similar…

机器学习 · 计算机科学 2021-10-12 Tarek Naous , Srinjay Sarkar , Abubakar Abid , James Zou

We present vivid, an R package for visualizing variable importance and variable interactions in machine learning models. The package provides a range of displays including heatmap and graph-based displays for viewing variable importance and…

统计计算 · 统计学 2024-10-15 Alan Inglis , Andrew Parnell , Catherine Hurley

It is critical to accurately simulate data when employing Monte Carlo techniques and evaluating statistical methodology. Measurements are often correlated and high dimensional in this era of big data, such as data obtained in…

Epidemiology aims at identifying subpopulations of cohort participants that share common characteristics (e.g. alcohol consumption) to explain risk factors of diseases in cohort study data. These data contain information about the…

We present a visual analytics system for exploring group differences in tensor fields with respect to all six degrees of freedom that are inherent in symmetric second-order tensors. Our framework closely integrates quantitative analysis,…

The causalimages R package enables causal inference with image and image sequence data, providing new tools for integrating novel data sources like satellite and bio-medical imagery into the study of cause and effect. One set of functions…

机器学习 · 计算机科学 2023-11-13 Connor T. Jerzak , Adel Daoud

Clustering in high-dimensions poses many statistical challenges. While traditional distance-based clustering methods are computationally feasible, they lack probabilistic interpretation and rely on heuristics for estimation of the number of…

统计方法学 · 统计学 2023-04-04 Abhinav Natarajan , Maria De Iorio , Andreas Heinecke , Emanuel Mayer , Simon Glenn