English
Related papers

Related papers: Characterising harmful data sources when construct…

200 papers

Large language models produce human-like text that drive a growing number of applications. However, recent literature and, increasingly, real world observations, have demonstrated that these models can generate language that is toxic,…

We propose a model-agnostic approach for mitigating the prediction bias of a black-box decision-maker, and in particular, a human decision-maker. Our method detects in the feature space where the black-box decision-maker is biased and…

Machine Learning · Computer Science 2020-11-18 Tong Wang , Maytal Saar-Tsechansky

High-speed flight vehicles, which travel much faster than the speed of sound, are crucial for national defense and space exploration. However, accurately predicting their behavior under numerous, varied flight conditions is a challenge and…

Machine Learning · Computer Science 2024-11-07 Tyler E. Korenyi-Both , Nathan J. Falkiewicz , Matthew C. Jones

When trained on diverse labeled data, machine learning models have proven themselves to be a powerful tool in all facets of society. However, due to budget limitations, deliberate or non-deliberate censorship, and other problems during data…

Machine Learning · Statistics 2022-03-25 Thomas Kehrenberg , Myles Bartlett , Viktoriia Sharmanska , Novi Quadrianto

This paper introduces a novel two-stage machine learning-based surrogate modeling framework to address inverse problems in scientific and engineering fields. In the first stage of the proposed framework, a machine learning model termed the…

Machine Learning · Computer Science 2024-01-05 Farhad Pourkamali-Anaraki , Jamal F. Husseini , Evan J. Pineda , Brett A. Bednarcyk , Scott E. Stapleton

In statistics, it is important to have realistic data sets available for a particular context to allow an appropriate and objective method comparison. For many use cases, benchmark data sets for method comparison are already available…

Applications · Statistics 2024-05-30 Maria Thurow , Ina Dormuth , Christina Sauer , Marc Ditzhaus , Markus Pauly

The surrogate data method is widely applied as a data dependent technique to test observed time series against a barrage of hypotheses. However, often the hypotheses one is able to address are not those of greatest interest, particularly…

Chaotic Dynamics · Physics 2007-05-23 Xiaodong Luo , Tomomichi Nakamura , Michael Small

The proliferation of Internet memes in the age of social media necessitates effective identification of harmful ones. Due to the dynamic nature of memes, existing data-driven models may struggle in low-resource scenarios where only a few…

Computation and Language · Computer Science 2024-11-11 Jianzhao Huang , Hongzhan Lin , Ziyan Liu , Ziyang Luo , Guang Chen , Jing Ma

This paper considers the surrogate modeling of a complex numerical code in a multifidelity framework when the code output is a time series. Using an experimental design of the low-and high-fidelity code levels, an original Gaussian process…

Statistics Theory · Mathematics 2022-02-24 Baptiste Kerleguer

In real case applications within the virtual prototyping process, it is not always possible to reduce the complexity of the physical models and to obtain numerical models which can be solved quickly. Usually, every single numerical…

Methodology · Statistics 2024-08-08 Thomas Most , Johannes Will

Surrogate algorithms such as Bayesian optimisation are especially designed for black-box optimisation problems with expensive objectives, such as hyperparameter tuning or simulation-based optimisation. In the literature, these algorithms…

Machine Learning · Computer Science 2024-03-14 Laurens Bliek , Arthur Guijt , Rickard Karlsson , Sicco Verwer , Mathijs de Weerdt

Data assimilation presents computational challenges because many high-fidelity models must be simulated. Various deep-learning-based surrogate modeling techniques have been developed to reduce the simulation costs associated with these…

Machine Learning · Computer Science 2022-12-28 Su Jiang , Louis J. Durlofsky

Investigators often use multi-source data (e.g., multi-center trials, meta-analyses of randomized trials, pooled analyses of observational cohorts) to learn about the effects of interventions in subgroups of some well-defined target…

Methodology · Statistics 2024-02-06 Guanbo Wang , Alexander Levis , Jon Steingrimsson , Issa Dahabreh

Despite the recent trend of developing and applying neural source code models to software engineering tasks, the quality of such models is insufficient for real-world use. This is because there could be noise in the source code corpora used…

Software Engineering · Computer Science 2022-10-04 Anh T. V. Dau , Thang Nguyen-Duc , Hoang Thanh-Tung , Nghi D. Q. Bui

Misleading or unnecessary data can have out-sized impacts on the health or accuracy of Machine Learning (ML) models. We present a Bayesian sequential selection method, akin to Bayesian experimental design, that identifies critically…

Machine Learning · Computer Science 2024-07-09 Ethan Pickering , Themistoklis P. Sapsis

Recently, several researchers have found that cost-based satisficing search with A* often runs into problems. Although some "work arounds" have been proposed to ameliorate the problem, there has been little concerted effort to pinpoint its…

Artificial Intelligence · Computer Science 2014-11-04 William Cushing , J. Benton , Patrick Eyerich , Subbarao Kambhampati

Surrogate markers are often employed in clinical trials to replace primary outcomes that may be difficult, expensive, or time-consuming to measure directly. These markers can accelerate the evaluation of new treatments, provided they…

Methodology · Statistics 2026-03-17 Pietro Carlotti , Layla Parast

Evaluating generative models for synthetic medical imaging is crucial yet challenging, especially given the high standards of fidelity, anatomical accuracy, and safety required for clinical applications. Standard evaluation of generated…

Image and Video Processing · Electrical Eng. & Systems 2025-05-13 Yash Deo , Yan Jia , Toni Lassila , William A. P. Smith , Tom Lawton , Siyuan Kang , Alejandro F. Frangi , Ibrahim Habli

Surrogate endpoints play an important role in drug development when they can be used to measure treatment effect early compared to the final clinical outcome and to predict clinical benefit or harm. Such endpoints are assessed for their…

Experts classifying data are often imprecise. Recently, several models have been proposed to train classifiers using the noisy labels generated by these experts. How to choose between these models? In such situations, the true labels are…

Methodology · Statistics 2014-05-15 Rafael Izbicki , Rafael Bassi Stern
‹ Prev 1 3 4 5 6 7 10 Next ›