English
Related papers

Related papers: A General Simulation-Based Optimisation Framework …

200 papers

Augmenting randomized controlled trials (RCTs) with external real-world data (RWD) has the potential to improve the finite sample efficiency of treatment effect estimators. We describe using adaptive targeted maximum likelihood estimation…

Methodology · Statistics 2025-01-30 Sky Qiu , Jens Tarp , Andrew Mertens , Mark van der Laan

Test-time optimization remains impractical at scale due to prohibitive inference costs--techniques like iterative refinement and multi-step verification can require $10-100\times$ more compute per query than standard decoding. Latent space…

Machine Learning · Computer Science 2025-11-10 Nathan Egbuna , Saatvik Gaur , Sunishchal Dev , Ashwinee Panda , Maheep Chaudhary

Testing and evaluation is an important step before the large-scale application of the autonomous driving systems (ADSs). Based on the three level of scenario abstraction theory, a testing can be performed within a logical scenario, followed…

Artificial Intelligence · Computer Science 2025-10-24 Xinzheng Wu , Junyi Chen , Jianfeng Wu , Longgao Zhang , Tian Xia , Yong Shen

Stress testing is an approach for evaluating the reliability of systems under extreme conditions which help reveal vulnerable scenarios that standard testing may overlook. Identifying such scenarios is of great importance in autonomous…

Robotics · Computer Science 2024-09-20 Linh Trinh , Quang-Hung Luu , Thai M. Nguyen , Hai L. Vu

Many modern products exhibit high reliability under normal operating conditions. Conducting life tests under these conditions may result in very few observed failures, insufficient for accurate inferences. Instead, accelerated life tests…

Applications · Statistics 2024-09-25 Narayanaswamy Balakrishnan , María Jaenada , Leandro Pardo

Many highly reliable products are designed to function for years without failure. For such systems accelerated degradation testing may provide significance information about the reliability properties of the system. In this paper, we…

Applications · Statistics 2021-10-13 Helmi Shat , Rainer Schwabe

Traditional benchmarks for large language models (LLMs), such as HELM and AIR-BENCH, primarily assess safety through breadth-oriented evaluation across diverse tasks and risk categories. However, real-world deployment often exposes a…

Machine Learning · Computer Science 2026-04-29 Keita Broadwater

The rapid proliferation of large language models (LLMs) in healthcare creates an urgent need for scalable and psychometrically sound evaluation methods. Conventional static benchmarks are costly to administer repeatedly, vulnerable to data…

Computation and Language · Computer Science 2026-03-26 Tianpeng Zheng , Zhehan Jiang , Jiayi Liu , Shicong Feng

The sequential multiple assignment randomized trial (SMART) is the ideal study design for the evaluation of multistage treatment regimes, which comprise sequential decision rules that recommend treatments for a patient at each of a series…

Methodology · Statistics 2024-05-15 Anastasios A. Tsiatis , Marie Davidian

The effectiveness of univariate forecasting models is often hampered by conditions that cause them stress. A model is considered to be under stress if it shows a negative behaviour, such as higher-than-usual errors or increased uncertainty.…

Machine Learning · Computer Science 2024-07-03 Ricardo Inácio , Vitor Cerqueira , Marília Barandas , Carlos Soares

This study combines simulated annealing with delta evaluation to solve the joint stratification and sample allocation problem. In this problem, atomic strata are partitioned into mutually exclusive and collectively exhaustive strata. Each…

Artificial Intelligence · Computer Science 2021-11-23 Mervyn O'Luing , Steven Prestwich , S. Armagan Tarim

Validating the safety of autonomous systems generally requires the use of high-fidelity simulators that adequately capture the variability of real-world scenarios. However, it is generally not feasible to exhaustively search the space of…

Machine Learning · Computer Science 2021-07-28 Mark Koren , Ahmed Nassar , Mykel J. Kochenderfer

Test-Time Optimization enables models to adapt to new data during inference by updating parameters on-the-fly. Recent advances in Vision-Language Models (VLMs) have explored learning prompts at test time to improve performance in downstream…

Computer Vision and Pattern Recognition · Computer Science 2025-07-29 Dhruv Sarkar , Aprameyo Chakrabartty , Bibhudatta Bhanja

What can be considered an appropriate statistical method for the primary analysis of a randomized clinical trial (RCT) with a time-to-event endpoint when we anticipate non-proportional hazards owing to a delayed effect? This question has…

Methodology · Statistics 2023-04-18 José L. Jiménez , Isobel Barrott , Francesca Gasperoni , Dominic Magirr

Traditional step-stress accelerated life testing models assume that test units originate from a homogeneous population. Recently, Lu and Kateri (2025) proposed a heterogeneous cumulative exposure based SSALT model to account for the…

Methodology · Statistics 2026-05-21 Pranoy Palit , Ayan Pal , Kiran Prajapat

We introduce adaptive learn-then-test (aLTT), an efficient hyperparameter selection procedure that provides finite-sample statistical guarantees on the population risk of AI models. Unlike the existing learn-then-test (LTT) technique, which…

Machine Learning · Statistics 2025-02-03 Matteo Zecchin , Sangwoo Park , Osvaldo Simeone

Test-time adaptation (TTA) addresses distribution shifts for streaming test data in unsupervised settings. Currently, most TTA methods can only deal with minor shifts and rely heavily on heuristic and empirical studies. To advance TTA under…

Machine Learning · Computer Science 2024-04-09 Shurui Gui , Xiner Li , Shuiwang Ji

In this paper, we consider the problem of long tail scenario modeling with budget limitation, i.e., insufficient human resources for model training stage and limited time and computing resources for model inference stage. This problem is…

Machine Learning · Computer Science 2023-06-30 Ya-Lin Zhang , Jun Zhou , Yankun Ren , Yue Zhang , Xinxing Yang , Meng Li , Qitao Shi , Longfei Li

In recent times, products have become increasingly complex and highly reliable, so failures typically occur after long periods of operation under normal conditions and may arise from multiple causes. This paper employs simple step-stress…

Methodology · Statistics 2025-08-01 Rathin Das , Soumya Roy , Biswabrata Pradhan

During the development of autonomous systems such as driverless cars, it is important to characterize the scenarios that are most likely to result in failure. Adaptive Stress Testing (AST) provides a way to search for the most-likely…

Machine Learning · Computer Science 2019-07-17 Mark Koren , Mykel Kochenderfer