中文
相关论文

相关论文: A nonmanipulable test

200 篇论文

We formalize and analyze a new problem in formal language theory termed control improvisation. Given a specification language, the problem is to produce an improviser, a probabilistic algorithm that randomly generates words in the language,…

形式语言与自动机理论 · 计算机科学 2017-04-24 Daniel J. Fremont , Alexandre Donzé , Sanjit A. Seshia

Phase III randomized clinical trials play a monumentally critical role in the evaluation of new medical products. Because of the intrinsic nature of uncertainty embedded in our capability in assessing the efficacy of a medical product,…

统计方法学 · 统计学 2019-02-25 Changyu Shen , Xiaochun Li

Randomized trials are widely considered as the gold standard for evaluating the effects of decision policies. Trial data is, however, drawn from a population which may differ from the intended target population and this raises a problem of…

统计方法学 · 统计学 2024-10-30 Sofia Ek , Dave Zachariah

Retrospective testing of predictive models does not consider the real-world context in which models are deployed. Prospective validation, on the other hand, enables meaningful comparisons between data generation processes by incorporating…

机器学习 · 计算机科学 2020-11-19 Steven Kearnes

Manufacturers of safety-critical systems must make the case that their product is sufficiently safe for public deployment. Much of this case often relies upon critical event outcomes from real-world testing, requiring manufacturers to be…

人工智能 · 计算机科学 2018-05-22 Jeremy Morton , Tim A. Wheeler , Mykel J. Kochenderfer

Motivated by the practice of exploratory research, we formulate an approach to multiple testing that reverses the conventional roles of the user and the multiple testing procedure. Traditionally, the user chooses the error criterion, and…

统计方法学 · 统计学 2013-10-03 Jelle J. Goeman , Aldo Solari

The increasing adoption of machine learning tools has led to calls for accountability via model interpretability. But what does it mean for a machine learning model to be interpretable by humans, and how can this be assessed? We focus on…

机器学习 · 计算机科学 2019-08-06 Dylan Slack , Sorelle A. Friedler , Carlos Scheidegger , Chitradeep Dutta Roy

If AI is the new electricity, what should we do to keep ourselves from getting electrocuted? In this work, we explore factors related to the potential of large language models (LLMs) to manipulate human decisions. We describe the results of…

人机交互 · 计算机科学 2024-10-01 Piotr Wilczyński , Wiktoria Mieleszczenko-Kowszewicz , Przemysław Biecek

Agent-based models play an important role in simulating complex emergent phenomena and supporting critical decisions. In this context, a software fault may result in poorly informed decisions that lead to disastrous consequences. The…

软件工程 · 计算机科学 2021-03-19 Andrew G. Clark , Neil Walkinshaw , Robert M. Hierons

Classical tests for a difference in means control the type I error rate when the groups are defined a priori. However, when the groups are instead defined via clustering, then applying a classical test yields an extremely inflated type I…

统计方法学 · 统计学 2022-11-01 Lucy L. Gao , Jacob Bien , Daniela Witten

Many automated system analysis techniques (e.g., model checking, model-based testing) rely on first obtaining a model of the system under analysis. System modeling is often done manually, which is often considered as a hindrance to adopt…

软件工程 · 计算机科学 2019-11-22 Jingyi Wang , Jun Sun , Qixia Yuan , Jun Pang

A benefit of randomized experiments is that covariate distributions of treatment and control groups are balanced on average, resulting in simple unbiased estimators for treatment effects. However, it is possible that a particular…

统计方法学 · 统计学 2019-02-01 Zach Branson , Luke Miratrix

Testing is a core software development activity that has huge potential to make software development more sustainable. In this paper, we discuss how environmental, social, economic, and technical sustainability map onto the activities of…

软件工程 · 计算机科学 2021-04-06 Armin Beer , Michael Felderer , Tobias Lorey , Stefan Mohacsi

There is a useful counterpart of conformal prediction for e-values, called conformal e-prediction. Conformal prediction can serve as basis for testing the assumption of exchangeability, leading to conformal testing. Similarly, conformal…

统计理论 · 数学 2024-11-05 Vladimir Vovk , Ilia Nouretdinov , Alex Gammerman

Permutation tests are a powerful and flexible approach to inference via resampling. As computational methods become more ubiquitous in the statistics curriculum, use of permutation tests has become more tractable. At the heart of the…

统计方法学 · 统计学 2025-06-09 Johanna Hardin , Lauren Quesada , Julie Ye , Nicholas J. Horton

Practical employment of Bayesian trial designs is still rare. Even if accepted in principle, the regulators have commonly required that such designs be calibrated according to an upper bound for the frequentist type I error rate. This…

统计方法学 · 统计学 2026-03-25 Elja Arjas , Dario Gasbarra

The Gibbard-Satterthwaite theorem states that no unanimous and non-dictatorial voting rule is strategyproof. We revisit voting rules and consider a weaker notion of strategyproofness called not obvious manipulability that was proposed by…

计算机科学与博弈论 · 计算机科学 2022-06-15 Haris Aziz , Alexander Lam

We analyze theoretical properties of the hybrid test for superior predictability. We demonstrate with a simple example that the test may not be pointwise asymptotically of level $\alpha$ at commonly used significance levels and may lead to…

计量经济学 · 经济学 2021-09-13 Deborah Kim

Under a regularity assumption we prove that reachability in fixed time for nonlinear control systems is robust under control sampling.

最优化与控制 · 数学 2020-06-22 Loïc Bourdin , Emmanuel Trélat

Random testing approaches work by generating inputs at random, or by selecting inputs randomly from some pre-defined operational profile. One long-standing question that arises in this and other testing contexts is as follows: When can we…

软件工程 · 计算机科学 2024-06-25 Neil Walkinshaw , Michael Foster , Jose Miguel Rojas , Robert M Hierons
‹ 上一页 1 8 9 10 下一页 ›