English
Related papers

Related papers: The notion of validity in experimental crowd dynam…

200 papers

Online labor markets have great potential as platforms for conducting experiments, as they provide immediate access to a large and diverse subject pool and allow researchers to conduct randomized controlled trials. We argue that online…

Human-Computer Interaction · Computer Science 2025-03-19 John J. Horton , David G. Rand , Richard J. Zeckhauser

Development of several alternative mathematical models for the biological system in question and discrimination between such models using experimental data is the best way to robust conclusions. Models which challenge existing theories are…

Quantitative Methods · Quantitative Biology 2016-02-01 Vitaly V. Ganusov

Multiple testing problems arise naturally in scientific studies because of the need to capture or convey more information with more variables. The literature is enormous, but the emphasis is primarily methodological, providing numerous…

Other Statistics · Statistics 2020-10-07 Yudi Pawitan , Arvid Sjölander

In recent years, human behavior simulation has drawn increasing attention from both academia and industry. The reasons fall into two aspects. First, simulation serves as a critical tool for understanding human behaviors, which has become…

Human-Computer Interaction · Computer Science 2024-12-12 Zhang Guozhen , Yu Zihan , Li Nian , Yu Fudan , Long Qingyue , Jin Depeng , Li Yong

We consider the problem of estimating personalized treatment policies that are "externally valid" or "generalizable": they perform well in target populations that differ from the experimental (or training) population from which the data are…

Econometrics · Economics 2025-11-10 Christopher Adjaho , Timothy Christensen

Searching for clues, gathering evidence, and reviewing case files are all techniques used by criminal investigators to draw sound conclusions and avoid wrongful convictions. Similarly, in software engineering (SE) research, we can develop…

Software Engineering · Computer Science 2023-09-19 Marvin Muñoz Barón , Marvin Wyrich , Daniel Graziotin , Stefan Wagner

Researchers must ensure that the claims about the knowledge produced by their work are valid. However, validity is neither well-understood nor consistently established in design science, which involves the development and evaluation of…

Software Engineering · Computer Science 2025-04-02 K. Larsen , R. Lukyanenko , Roland M. Mueller , V. Storey , J. Parsons , D. Vandermeer , D. Hovorka

There are many cluster analysis methods that can produce quite different clusterings on the same dataset. Cluster validation is about the evaluation of the quality of a clustering; "relative cluster validation" is about using such criteria…

Methodology · Statistics 2020-09-10 Christian Hennig

Policy decisions often depend on evidence generated elsewhere. We take a Bayesian decision-theoretic approach to choosing where to experiment to optimize external validity. We frame external validity through a policy lens, developing a…

When assessing causal effects, determining the target population to which the results are intended to generalize is a critical decision. Randomized and observational studies each have strengths and limitations for estimating causal effects…

Methodology · Statistics 2022-10-21 Irina Degtiar , Sherri Rose

There is a well-known problem in Null Hypothesis Significance Testing: many statistically significant results fail to replicate in subsequent experiments. We show that this problem arises because standard `point-form null' significance…

Methodology · Statistics 2025-02-06 Fintan Costello , Paul Watts

Systematic differences in experimental materials, methods, measurements, and data handling between labs, over time, and among personnel can sabotage experimental reproducibility. Uncovering such differences can be difficult and time…

Applications · Statistics 2018-07-18 Bert Gunter

Evidence shows that in a significant number of cases the current methods of research do not allow for reproducible and falsifiable procedures of scientific investigation. As a consequence, the majority of critical decisions at all levels,…

Computational Finance · Quantitative Finance 2018-09-20 Jorge Faleiro

Measuring interdisciplinarity is a pertinent but challenging issue in quantitative studies of science. There seems to be a consensus in the literature that the concept of interdisciplinarity is multifaceted and ambiguous. Unsurprisingly,…

Digital Libraries · Computer Science 2019-12-11 Qi Wang , Jesper Wiborg Schneider

In physics the value of a theory is measured by its agreement with experimental data. But how should the physics community gauge the value of an emerging theory that has not been tested experimentally as of yet? With no reality check, a…

Instrumentation and Methods for Astrophysics · Physics 2011-08-29 Abraham Loeb

We proposed a probabilistic approach to joint modeling of participants' reliability and humans' regularity in crowdsourced affective studies. Reliability measures how likely a subject will respond to a question seriously; and regularity…

Machine Learning · Statistics 2017-01-09 Jianbo Ye , Jia Li , Michelle G. Newman , Reginald B. Adams , James Z. Wang

We present a strategy capable of describing basic features of the dynamics of crowds. The behaviour of the crowd is considered from a twofold perspective. We examine both the large scale behaviour of the crowd, and phenomena happening at…

Mathematical Physics · Physics 2011-08-09 Joep Evers

Relevance is an underlying concept in the field of Information Science and Retrieval. It is a cognitive notion consisting of several different criteria or dimensions. Theoretical models of relevance allude to interdependence between these…

Information Retrieval · Computer Science 2019-07-26 Sagar Uprety , Shahram Dehdashti , Lauren Fell , Peter Bruza , Dawei Song

In multiple hypothesis testing, the volume of data, defined as the number of replications per null times the total number of nulls, usually defines the amount of resource required. On the other hand, power is an important measure of…

Statistics Theory · Mathematics 2009-06-05 Zhiyi Chi

Large language models (LLMs) are increasingly used as decision-support tools in data-constrained scientific workflows, where correctness and validity are critical. However, evaluation practices often emphasize stability or reproducibility…

Machine Learning · Computer Science 2026-03-18 Nazia Riasat