中文
相关论文

相关论文: Characterising harmful data sources when construct…

200 篇论文

Accurately estimating model performance poses a significant challenge, particularly in scenarios where the source and target domains follow different data distributions. Most existing performance prediction methods heavily rely on the…

计算机视觉与模式识别 · 计算机科学 2024-08-07 Ekaterina Khramtsova , Mahsa Baktashmotlagh , Guido Zuccon , Xi Wang , Mathieu Salzmann

Aerodynamic shape optimization in industry still faces challenges related to robustness and scalability. This aspect becomes crucial for advanced optimizations that rely on expensive high-fidelity flow solvers, where computational budget…

流体动力学 · 物理学 2025-05-26 Marc Schouler , Anca Belme , Paola Cinnella

Machine learning methods are increasingly used to build computationally inexpensive surrogates for complex physical models. The predictive capability of these surrogates suffers when data are noisy, sparse, or time-dependent. As we are…

机器学习 · 计算机科学 2024-05-20 A. Diaw , M. McKerns , I. Sagert , L. G. Stanton , M. S. Murillo

Engineers widely use Gaussian process regression framework to construct surrogate models aimed to replace computationally expensive physical models while exploring design space. Thanks to Gaussian process properties we can use both samples…

机器学习 · 统计学 2017-07-14 Evgeny Burnaev , Alexey Zaytsev

Complex black-box predictive models may have high accuracy, but opacity causes problems like lack of trust, lack of stability, sensitivity to concept drift. On the other hand, interpretable models require more work related to feature…

机器学习 · 计算机科学 2019-03-01 Alicja Gosiewska , Aleksandra Gacek , Piotr Lubon , Przemyslaw Biecek

Systematic exploration of Agent-Based Models (ABMs) is challenged by the curse of dimensionality and their inherent stochasticity. We present a multi-stage pipeline integrating the systematic design of experiments with machine learning…

机器学习 · 计算机科学 2026-04-07 Paul Saves , Matthieu Mastio , Nicolas Verstaevel , Benoit Gaudou

Recent years have experienced increasing utilization of complex machine learning models across multiple sources of data to inform more generalizable decision-making. However, distribution shifts across data sources and privacy concerns…

统计方法学 · 统计学 2024-05-16 Yi Liu , Alexander W. Levis , Sharon-Lise Normand , Larry Han

Bayesian regression determines model parameters by minimizing the expected loss, an upper bound to the true generalization error. However, the loss ignores misspecification, where models are imperfect. Parameter uncertainties from Bayesian…

机器学习 · 统计学 2024-11-07 Thomas D Swinburne , Danny Perez

A machine-learning-based framework for modeling the error introduced by surrogate models of parameterized dynamical systems is proposed. The framework entails the use of high-dimensional regression techniques (e.g., random forests, LASSO)…

数值分析 · 计算机科学 2017-06-02 Sumeet Trehan , Kevin Carlberg , Louis J. Durlofsky

Data-driven modeling can suffer from a constant demand for data, leading to reduced accuracy and impractical for engineering applications due to the high cost and scarcity of information. To address this challenge, we propose a progressive…

机器学习 · 计算机科学 2023-10-09 Teeratorn Kadeethum , Daniel O'Malley , Youngsoo Choi , Hari S. Viswanathan , Hongkyu Yoon

One of the main challenges in surrogate modeling is the limited availability of data due to resource constraints associated with computationally expensive simulations. Multi-fidelity methods provide a solution by chaining models in a…

Among adversarial attacks against sequential recommender systems, model extraction attacks represent a method to attack sequential recommendation models without prior knowledge. Existing research has primarily concentrated on the…

机器学习 · 计算机科学 2026-03-04 Hui Zhang , Fu Liu

Fatigue crack growth is one of the most common types of deterioration in metal structures with significant implications on their reliability. Recent advances in Structural Health Monitoring (SHM) have motivated the use of structural…

机器学习 · 统计学 2023-10-12 Nicholas E. Silionis , Konstantinos N. Anyfantis

Motivated by the desire to generate labels for real-time data we develop a method to estimate the dependency structure and accuracy of weak supervision sources incrementally. Our method first estimates the dependency structure associated…

机器学习 · 计算机科学 2022-05-12 Richard Gresham Correro

In engineering design and scientific computing, computational cost and predictive accuracy are intrinsically coupled. High-fidelity simulations provide accurate predictions but at substantial computational costs, while lower-fidelity…

机器学习 · 计算机科学 2026-05-11 Ahmed Mohamed Eisa Nasr , Ali Elham , Haris Moazam Sheikh

Adaptive designs are increasingly used in clinical trials and online experiments to improve participant outcomes by dynamically updating treatment allocation as data accumulate. In practice, experimenters often consider multiple candidate…

统计方法学 · 统计学 2026-04-08 Wenxin Zhang , Aaron Hudson , Maya Petersen , Mark van der Laan

Measurement-constrained datasets, often encountered in semi-supervised learning, arise when data labeling is costly, time-intensive, or hindered by confidentiality or ethical concerns, resulting in a scarcity of labeled data. In certain…

统计方法学 · 统计学 2025-01-15 Yixin Shen , Yang Ning

Surrogate model-based optimization has been increasingly used in the field of engineering design. It involves creating a surrogate model with objective functions or constraints based on the data obtained from simulations or real-world…

最优化与控制 · 数学 2023-08-28 Minyoung Jwa , Jihoon Kim , Seungyeon Shin , Ah-hyeon Jin , Dongju Shin , Namwoo Kang

Labeling training data is a key bottleneck in the modern machine learning pipeline. Recent weak supervision approaches combine labels from multiple noisy sources by estimating their accuracies without access to ground truth labels; however,…

机器学习 · 统计学 2019-03-15 Paroma Varma , Frederic Sala , Ann He , Alexander Ratner , Christopher Ré

Given the long follow-up periods that are often required for treatment or intervention studies, the potential to use surrogate markers to decrease the required follow-up time is a very attractive goal. However, previous studies have shown…

统计方法学 · 统计学 2016-08-12 Layla Parast , Tianxi Cai , Lu Tian