中文
相关论文

相关论文: Extreme Value Distribution Based Gene Selection Cr…

200 篇论文

In the covariate shift learning scenario, the training and test covariate distributions differ, so that a predictor's average loss over the training and test distributions also differ. In this work, we explore the potential of extreme…

机器学习 · 计算机科学 2018-03-13 Fulton Wang , Cynthia Rudin

Machine Learning methods have of late made significant efforts to solving multidisciplinary problems in the field of cancer classification using microarray gene expression data. Feature subset selection methods can play an important role in…

计算工程、金融与科学 · 计算机科学 2013-03-04 G. Prat , Ll. Belanche

A key obstacle in automated analytics and meta-learning is the inability to recognize when different datasets contain measurements of the same variable. Because provided attribute labels are often uninformative in practice, this task may be…

机器学习 · 计算机科学 2019-09-12 Jonas Mueller , Alex Smola

Algorithms for machine learning-guided design, or design algorithms, use machine learning-based predictions to propose novel objects with desired property values. Given a new design task -- for example, to design novel proteins with high…

机器学习 · 计算机科学 2025-07-04 Clara Fannjiang , Ji Won Park

For measuring tail risk with scarce extreme events, extreme value analysis is often invoked as the statistical tool to extrapolate to the tail of a distribution. The presence of large datasets benefits tail risk analysis by providing more…

统计方法学 · 统计学 2023-12-18 Liujun Chen , Deyuan Li , Chen Zhou

We study distributional robustness in the context of Extreme Value Theory (EVT). We provide a data-driven method for estimating extreme quantiles in a manner that is robust against incorrect model assumptions underlying the application of…

统计理论 · 数学 2020-06-09 Jose Blanchet , Fei He , Karthyek R. A. Murthy

The problem of selecting the most useful features from a great many (eg, thousands) of candidates arises in many areas of modern sciences. An interesting problem from genomic research is that, from thousands of genes that are active…

应用统计 · 统计学 2018-05-15 Longhai Li , Weixin Yao

The problem of overdispersion in multivariate count data is a challenging issue. Nowadays, it covers a central role mainly due to the relevance of modern technologies data, such as Next Generation Sequencing and textual data from the web or…

统计方法学 · 统计学 2025-02-24 Noemi Corsini , Cinzia Viroli

Logistic regression is a common classification method in supervised learning. Surprisingly, there are very few solutions for performing logistic regression with missing values in the covariates. We suggest a complete approach based on a…

统计方法学 · 统计学 2019-08-09 Wei Jiang , Julie Josse , Marc Lavielle , TraumaBase Group

The r largest order statistics approach is widely used in extreme value analysis because it may use more information from the data than just the block maxima. In practice, the choice of r is critical. If r is too large, bias can occur; if…

统计方法学 · 统计学 2018-06-13 Brian Bader , Jun Yan , Xuebin Zhang

Motivation: Microarray data has been recently been shown to be efficacious in distinguishing closely related cell types that often appear in the diagnosis of cancer. It is useful to determine the minimum number of genes needed to do such a…

生物物理 · 物理学 2007-05-23 J. M. Deutsch

While deep generative models~(DGMs) have demonstrated remarkable success in capturing complex data distributions, they consistently fail to learn constraints that encode domain knowledge and thus require constraint integration. Existing…

机器学习 · 计算机科学 2025-02-13 Ruoyan Li , Dipti Ranjan Sahu , Guy Van den Broeck , Zhe Zeng

Missing values are largely inevitable in gene expression microarray studies. Data sets often have significant omissions due to individuals dropping out of experiments, errors in data collection, image corruptions, and so on. Missing data…

定量方法 · 定量生物学 2018-09-18 Marie Li

We propose a novel regression adjustment method designed for estimating distributional treatment effect parameters in randomized experiments. Randomized experiments have been extensively used to estimate treatment effects in various…

计量经济学 · 经济学 2024-07-24 Undral Byambadalai , Tatsushi Oka , Shota Yasui

Classification in the dissimilarity space has become a very active research area since it provides a possibility to learn from data given in the form of pairwise non-metric dissimilarities, which otherwise would be difficult to cope with.…

Logistic regression is among the most widely used statistical methods for linear discriminant analysis. In many applications, we only observe possibly mislabeled responses. Fitting a conventional logistic regression can then lead to biased…

应用统计 · 统计学 2017-02-21 Hung Hung , Zhi-Yu Jou , Su-Yun Huang

Data-driven anomaly detection methods typically build a model for the normal behavior of the target system, and score each data instance with respect to this model. A threshold is invariably needed to identify data instances with high (or…

机器学习 · 统计学 2019-10-09 Sreelekha Guggilam , S. M. Arshad Zaidi , Varun Chandola , Abani Patra

Selective prediction, where a model has the option to abstain from making a decision, is crucial for machine learning applications in which mistakes are costly. In this work, we focus on distributional regression and introduce a framework…

统计理论 · 数学 2025-04-01 Ahmed Zaoui , Clément Dombry

We propose a modification of linear discriminant analysis, referred to as compressive regularized discriminant analysis (CRDA), for analysis of high-dimensional datasets. CRDA is specially designed for feature elimination purpose and can be…

统计方法学 · 统计学 2018-04-12 Muhammad Naveed Tabassum , Esa Ollila

Statistical analysis of social networks provides valuable insights into complex network interactions across various scientific disciplines. However, accurate modeling of networks remains challenging due to the heavy computational burden and…

社会与信息网络 · 计算机科学 2023-07-25 Helal El-Zaatari , Fei Yu , Michael R Kosorok