English
Related papers

Related papers: Analysis in HUGIN of Data Conflict

200 papers

Mutual information is commonly used as a measure of similarity between competing labelings of a given set of objects, for example to quantify performance in classification and community detection tasks. As argued recently, however, the…

Social and Information Networks · Computer Science 2025-07-17 Maximilian Jerdee , Alec Kirkley , M. E. J. Newman

Classifiers are often tested on relatively small data sets, which should lead to uncertain performance metrics. Nevertheless, these metrics are usually taken at face value. We present an approach to quantify the uncertainty of…

Machine Learning · Statistics 2021-03-05 Niklas Tötsch , Daniel Hoffmann

We are able to unify various disparate claims and results in the literature, that stand in the way of a unified description and understanding of human conflict. First, we provide a reconciliation of the numerically different exponent values…

Physics and Society · Physics 2019-11-06 Michael Spagat , Stijn van Weezel , Minzhang Zheng , Neil F. Johnson

Unmeasured confounding is a major challenge for identifying causal relationships from non-experimental data. Here, we propose a method that can accommodate unmeasured discrete confounding. Extending recent identifiability results in deep…

Machine Learning · Computer Science 2024-08-13 Patrick Burauel , Frederick Eberhardt , Michel Besserve

Human disease diagnosis is a complicated process and requires high level of expertise. Any attempt of developing a web-based expert system dealing with human disease diagnosis has to overcome various difficulties. This paper describes a…

Artificial Intelligence · Computer Science 2010-06-24 Mir Anamul Hasan , Khaja Md. Sher-E-Alam , Ahsan Raja Chowdhury

In many areas of science multiple sets of data are collected pertaining to the same system. Examples are food products which are characterized by different sets of variables, bio-processes which are on-line sampled with different…

A growing body of literature attempts to learn about contagion using observational (i.e. non-experimental) data collected from a single social network. While the conclusions of these studies may be correct, the methods rely on assumptions…

Applications · Statistics 2017-06-30 Elizabeth L. Ogburn

Due to the complexity of the human body, most diseases present a high inter-personal variability in the way they manifest, i.e. in their phenotype, which has important clinical repercussions - as for instance the difficulty in defining…

Physics and Society · Physics 2018-06-06 Massimiliano Zanin , Juan Manuel Tuñas , Ernestina Menasalvas

We review possible measures of complexity which might in particular be applicable to situations where the complexity seems to arise spontaneously. We point out that not all of them correspond to the intuitive (or "naive") notion, and that…

Data Analysis, Statistics and Probability · Physics 2012-08-20 Peter Grassberger

In a multi-source environment, each source has its own credibility. If there is no external knowledge about credibility then we can use the information provided by the sources to assess their credibility. In this paper, we propose a way to…

Signal Processing · Electrical Eng. & Systems 2018-03-14 Pan Wei , John E. Ball , Derek T. Anderson , Archit Harsh , Christopher Archibald

In this discussion we consider why it is important to estimate causal effect parameters well even they are not identified, propose a partially identified approach for causal inference in the presence of colliders, point out an…

Methodology · Statistics 2017-11-01 Edward H. Kennedy , Sivaraman Balakrishnan

With the growing interest in social applications of Natural Language Processing and Computational Argumentation, a natural question is how controversial a given concept is. Prior works relied on Wikipedia's metadata and on content analysis…

Computation and Language · Computer Science 2019-08-21 Benjamin Sznajder , Ariel Gera , Yonatan Bilu , Dafna Sheinwald , Ella Rabinovich , Ranit Aharonov , David Konopnicki , Noam Slonim

Traditionally, statistical and causal inference on human subjects rely on the assumption that individuals are independently affected by treatments or exposures. However, recently there has been increasing interest in settings, such as…

Methodology · Statistics 2020-02-25 Elizabeth L. Ogburn , Ilya Shpitser , Youjin Lee

Estimating how a treatment affects different individuals, known as heterogeneous treatment effect estimation, is an important problem in empirical sciences. In the last few years, there has been a considerable interest in adapting machine…

Machine Learning · Computer Science 2024-10-18 Christopher Tran , Keith Burghardt , Kristina Lerman , Elena Zheleva

Comparing two population means of network data is of paramount importance in a wide range of scientific applications. Many existing network inference solutions focus on global testing of entire networks, without comparing individual network…

Methodology · Statistics 2019-10-10 Yin Xia , Lexin Li

Knowledge about existence, strength, and dominant direction of causal influences is of paramount importance for understanding complex systems. With limited amounts of realistic data, however, current methods for investigating causal links…

Data Analysis, Statistics and Probability · Physics 2020-10-20 Erik Laminski , Klaus R. Pawelzik

Measurement involves the determination of quantitative estimates of physical quantities from experiment, along with estimates of their associated uncertainties. Herewith an experimental system model is the key to extracting information from…

Applications · Statistics 2008-09-01 Vladimir B. Bokov

Observational data is increasingly used as a means for making individual-level causal predictions and intervention recommendations. The foremost challenge of causal inference from observational data is hidden confounding, whose presence…

Machine Learning · Statistics 2018-10-30 Nathan Kallus , Aahlad Manas Puli , Uri Shalit

Large-scale data are often characterized by some degree of inhomogeneity as data are either recorded in different time regimes or taken from multiple sources. We look at regression models and the effect of randomly changing coefficients,…

Methodology · Statistics 2016-08-11 Nicolai Meinshausen , Peter Bühlmann

Today, data analysts largely rely on intuition to determine whether missing or withheld rows of a dataset significantly affect their analyses. We propose a framework that can produce automatic contingency analysis, i.e., the range of values…

Databases · Computer Science 2020-04-09 Xi Liang , Zechao Shang , Aaron J. Elmore , Sanjay Krishnan , Michael J. Franklin