中文
相关论文

相关论文: Finding Statistically Significant Attribute Intera…

200 篇论文

In this work we suggest a statistical mechanics approach to the classification of high-dimensional data according to a binary label. We propose an algorithm whose aim is twofold: First it learns a classifier from a relatively small number…

统计力学 · 物理学 2009-07-22 Andrea Pagnani , Francesca Tria , Martin Weigt

Discrimination discovery and prevention/removal are increasingly important tasks in data mining. Discrimination discovery aims to unveil discriminatory practices on the protected attribute (e.g., gender) by analyzing the dataset of…

机器学习 · 计算机科学 2016-11-23 Lu Zhang , Yongkai Wu , Xintao Wu

Hierarchically-organized data arise naturally in many psychology and neuroscience studies. As the standard assumption of independent and identically distributed samples does not hold for such data, two important problems are to accurately…

统计理论 · 数学 2018-09-03 Irene Dowding , Stefan Haufe

Machine learning models can automatically learn complex relationships, such as non-linear and interaction effects. Interpretable machine learning methods such as partial dependence plots visualize marginal feature effects but may lead to…

机器学习 · 统计学 2022-02-16 Julia Herbinger , Bernd Bischl , Giuseppe Casalicchio

Outlying observations are frequently encountered across a wide spectrum of scientific domains, posing notable challenges to the generalizability of statistical models and the reproducibility of downstream analysis. They are identified…

统计方法学 · 统计学 2026-03-17 Dongliang Zhang , Masoud Asgharian , Martin A. Lindquist

This paper reconsiders the problem of testing the equality of two unspecified continuous distributions. The framework, which we propose, allows for readable and insightful data visualisation and helps to understand and quantify how two…

统计方法学 · 统计学 2025-03-04 Bogdan Ćmiel , Teresa Ledwina

Training the deep neural networks that dominate NLP requires large datasets. These are often collected automatically or via crowdsourcing, and may exhibit systematic biases or annotation artifacts. By the latter we mean spurious…

计算与语言 · 计算机科学 2022-03-29 Pouya Pezeshkpour , Sarthak Jain , Sameer Singh , Byron C. Wallace

We introduce a new statistical test based on the observed spacings of ordered data. The statistic is sensitive to detect non-uniformity in random samples, or short-lived features in event time series. Under some conditions, this new test…

统计方法学 · 统计学 2022-10-27 Philipp Eller , Lolian Shtembari

For testing the statistical significance of a treatment effect, we usually compare between two parts of a population, one is exposed to the treatment, and the other is not exposed to it. Standard parametric and nonparametric two-sample…

统计计算 · 统计学 2012-11-02 Bikram Karmakar , Kumaresh Dhara , Kushal Kumar Dey , Analabha Basu , Anil Ghosh

The increasing availability of time --and space-- resolved data describing human activities and interactions gives insights into both static and dynamic properties of human behavior. In practice, nevertheless, real-world datasets can often…

A new type of statistical analysis of the science and technical information (STI) in the Web context is produced. We propose a set of indicators about Web users, visualized bibliographic records, and e-commercial transactions. In addition,…

信息检索 · 计算机科学 2008-11-06 Xavier Polanco , Ivana Roche , Dominique Besagni

Observed associations in a database may be due in whole or part to variations in unrecorded (latent) variables. Identifying such variables and their causal relationships with one another is a principal goal in many scientific and practical…

机器学习 · 计算机科学 2012-12-12 Ricardo Silva , Richard Scheines , Clark Glymour , Peter L. Spirtes

A trend in all scientific disciplines, based on advances in technology, is the increasing availability of high dimensional data in which are buried important information. A current urgent challenge to statisticians is to develop effective…

应用统计 · 统计学 2010-09-30 Herman Chernoff , Shaw-Hwa Lo , Tian Zheng

To date, testing interactions in high dimensions has been a challenging task. Existing methods often have issues with sensitivity to modeling assumptions and heavily asymptotic nominal p-values. To help alleviate these issues, we propose a…

机器学习 · 统计学 2012-06-29 Noah Simon , Robert Tibshirani

In medical organizations large amount of personal data are collected and analyzed by the data miner or researcher, for further perusal. However, the data collected may contain sensitive information such as specific disease of a patient and…

密码学与安全 · 计算机科学 2012-03-19 Pawan R Bhaladhare , Devesh Jinwala

This study introduces a nonparametric definition of interaction and provides an approach to both interaction discovery and efficient estimation of this parameter. Using stochastic shift interventions and ensemble machine learning, our…

统计方法学 · 统计学 2024-07-01 David B. McCoy , Alan E. Hubbard , Alejandro Schuler , Mark J. van der Laan

Instance-level image classification tasks have traditionally relied on single-instance labels to train models, e.g., few-shot learning and transfer learning. However, set-level coarse-grained labels that capture relationships among…

机器学习 · 计算机科学 2023-11-21 Renyu Zhang , Aly A. Khan , Yuxin Chen , Robert L. Grossman

PredDiff is a model-agnostic, local attribution method that is firmly rooted in probability theory. Its simple intuition is to measure prediction changes while marginalizing features. In this work, we clarify properties of PredDiff and its…

机器学习 · 计算机科学 2023-07-12 Stefan Blücher , Johanna Vielhaben , Nils Strodthoff

Accurately estimating human internal states, such as personality traits or behavioral patterns, is critical for enhancing the effectiveness of human-robot interaction, particularly in group settings. These insights are key in applications…

机器人学 · 计算机科学 2025-05-16 Xuebo Ji , Zherong Pan , Xifeng Gao , Lei Yang , Xinxin Du , Kaiyun Li , Yongjin Liu , Wenping Wang , Changhe Tu , Jia Pan

Rare properties remain a challenge for statistical model checking (SMC) due to the quadratic scaling of variance with rarity. We address this with a variance reduction framework based on lightweight importance splitting observers. These…

计算机科学中的逻辑 · 计算机科学 2015-04-29 Cyrille Jegourel , Axel Legay , Sean Sedwards , Louis-Marie Traonouez