中文
相关论文

相关论文: RLE Plots: Visualising Unwanted Variation in High …

200 篇论文

Non-linear dimensionality reduction can be performed by \textit{manifold learning} approaches, such as Stochastic Neighbour Embedding (SNE), Locally Linear Embedding (LLE) and Isometric Feature Mapping (ISOMAP). These methods aim to produce…

机器学习 · 统计学 2021-12-09 Theodoulos Rodosthenous , Vahid Shahrezaei , Marina Evangelou

Logs are an essential source of information for people to understand the running status of a software system. Due to the evolving modern software architecture and maintenance methods, more research efforts have been devoted to automated log…

软件工程 · 计算机科学 2024-04-09 Xingfang Wu , Heng Li , Foutse Khomh

Studying complex real-world phenomena often involves data from multiple views (e.g. sensor modalities or brain regions), each capturing different aspects of the underlying system. Within neuroscience, there is growing interest in…

High-dimensional data can often display heterogeneity due to heteroscedastic variance or inhomogeneous covariate effects. Penalized quantile and expectile regression methods offer useful tools to detect heteroscedasticity in…

统计方法学 · 统计学 2023-03-23 Rebeka Man , Kean Ming Tan , Zian Wang , Wen-Xin Zhou

Conformance checking techniques detect undesired process behavior by comparing process executions that are recorded in event logs to desired behavior that is captured in a dedicated process model. If such models are not available,…

机器学习 · 计算机科学 2025-05-29 Michael Grohs , Adrian Rebmann , Jana-Rebecca Rehse

Anomaly detection in visual data refers to the problem of differentiating abnormal appearances from normal cases. Supervised approaches have been successfully applied to different domains, but require an abundance of labeled data. Due to…

图像与视频处理 · 电气工程与系统科学 2021-04-29 Dejan Stepec , Danijel Skocaj

The spread-location plot has often been used as a diagnostic plot suitable for many types of fitted statistical models. The spread-location plot which plots the absolute residual or square-root absolute residual versus fitted value along…

统计理论 · 数学 2016-11-08 A. Ian McLeod

Positive-valued signal data is common in many biological and medical applications, where the data are often generated from imaging techniques such as mass spectrometry. In such a setting, the relative intensities of the raw features are…

统计方法学 · 统计学 2021-04-15 Stephen Bates , Robert Tibshirani

To solve key biomedical problems, experimentalists now routinely measure millions or billions of features (dimensions) per sample, with the hope that data science techniques will be able to build accurate data-driven inferences. Because…

Logistic regression is a ubiquitous method for probabilistic classification. However, the effectiveness of logistic regression depends upon careful and relatively computationally expensive tuning, especially for the regularisation…

机器学习 · 计算机科学 2025-04-04 Angus Dempster , Geoffrey I. Webb , Daniel F. Schmidt

Ordinary differential equations (ODEs) are widely used to model dynamical behavior of systems. It is important to perform identifiability analysis prior to estimating unknown parameters in ODEs (a.k.a. inverse problem), because if a system…

最优化与控制 · 数学 2021-03-11 Xing Qiu , Tao Xu , Babak Soltanalizadeh , Hulin Wu

Complex networks, modeled as large graphs, received much attention during these last years. However, data on such networks is only available through intricate measurement procedures. Until recently, most studies assumed that these…

网络与互联网体系结构 · 计算机科学 2007-05-23 Matthieu Latapy , Clemence Magnien

It is basic question in biology and other fields to identify the char- acteristic properties that on one hand are shared by structures from a particular realm, like gene regulation, protein-protein interaction or neu- ral networks or…

定量方法 · 定量生物学 2012-10-19 Anirban Banerjee , Jürgen Jost

Symbolic regression (SR) aims to discover explicit mathematical expressions that explain observed data and is widely used in domains where interpretability is essential. Because interpretability requires expressions to reflect meaningful…

神经与进化计算 · 计算机科学 2026-05-18 Koki Ikeda , Masahiro Nomura , Ryoki Hamano

We are living in the big data age: An ever increasing amount of data is being produced through data acquisition and computer simulations. While large scale analysis and simulations have received significant attention for cloud and…

图形学 · 计算机科学 2019-02-26 Stefan Eilemann

Matrix-valued data, where each observation is represented as a matrix, frequently arises in various scientific disciplines. Modeling such data often relies on matrix-variate normal distributions, making matrix-variate normality testing…

统计方法学 · 统计学 2025-05-02 Fen Jiang , Jianhua Zhao , Changchun Shang , Xuan Ma , Yue Wang , Ye Tao

Weird generalization is a phenomenon in which models fine-tuned on data from a narrow domain (e.g. insecure code) develop surprising traits that manifest even outside that domain (e.g. broad misalignment)-a phenomenon that prior work has…

计算与语言 · 计算机科学 2026-05-05 Miriam Wanner , Hannah Collison , William Jurayj , Benjamin Van Durme , Mark Dredze , William Walden

To gain insight into the mechanisms behind machine learning methods, it is crucial to establish connections among the features describing data points. However, these correlations often exhibit a high-dimensional and strongly nonlinear…

机器学习 · 计算机科学 2025-03-04 Lorenzo Basile , Santiago Acevedo , Luca Bortolussi , Fabio Anselmi , Alex Rodriguez

Analysis of high-dimensional data is currently a popular field of research, thanks to many applications e.g. in genetics (DNA data in genomewide association studies), spectrometry or web analysis. At the same time, the type of problems that…

统计方法学 · 统计学 2018-05-25 Jozef Jakubik

Big Data involves both a large number of events but also many variables. This paper will concentrate on the challenge presented by the large number of variables in a Big Dataset. It will start with a brief review of exploratory data…

应用统计 · 统计学 2019-07-24 S. J. Watts , L. Crow