中文
相关论文

相关论文: Effect sizes as a statistical feature-selector-bas…

200 篇论文

Penalized regression models such as the Lasso have proved useful for variable selection in many fields - especially for situations with high-dimensional data where the numbers of predictors far exceeds the number of observations. These…

统计方法学 · 统计学 2014-03-19 Kasper Brink-Jensen , Claus Thorn Ekstrøm

Educational data mining (EDM) is a new growing research area and the essence of data mining concepts are used in the educational field for the purpose of extracting useful information on the behaviors of students in the learning process. In…

数据库 · 计算机科学 2009-12-22 M. Ramaswami , R. Bhaskaran

Detecting dependence between variables is a crucial issue in statistical science. In this paper, we propose a novel metric called label projection correlation to measure the dependence between numerical and categorical variables. The…

统计方法学 · 统计学 2025-06-24 Yixiao Liu , Pengjian Shang

In machine learning, the exponential growth of data and the associated ``curse of dimensionality'' pose significant challenges, particularly with expansive yet sparse datasets. Addressing these challenges, multi-view ensemble learning (MEL)…

Few-shot learning is a relatively new technique that specializes in problems where we have little amounts of data. The goal of these methods is to classify categories that have not been seen before with just a handful of samples. Recent…

Eficient, physically-inspired descriptors of the structure and composition of molecules and materials play a key role in the application of machine-learning techniques to atomistic simulations. The proliferation of approaches, as well as…

计算物理 · 物理学 2020-12-11 Alexander Goscinski , Guillaume Fraux , Giulio Imbalzano , Michele Ceriotti

Feature selection is an important process in machine learning and knowledge discovery. By selecting the most informative features and eliminating irrelevant ones, the performance of learning algorithms can be improved and the extraction of…

机器学习 · 计算机科学 2024-01-17 Chunxu Cao , Qiang Zhang

This paper studies simultaneous feature selection and extraction in supervised and unsupervised learning. We propose and investigate selective reduced rank regression for constructing optimal explanatory factors from a parsimonious subset…

统计方法学 · 统计学 2016-10-27 Yiyuan She

Feature Learning aims to extract relevant information contained in data sets in an automated fashion. It is driving force behind the current deep learning trend, a set of methods that have had widespread empirical success. What is lacking…

机器学习 · 统计学 2015-04-02 Brendan van Rooyen , Robert C. Williamson

Mammography is the most effective and available tool for breast cancer screening. However, the low positive predictive value of breast biopsy resulting from mammogram interpretation leads to approximately 70% unnecessary biopsies with…

机器学习 · 计算机科学 2013-06-04 Sahar A. Mokhtar , Alaa. M. Elsayad

Recent discussion of the success of feature selection methods has argued that focusing on a relatively small number of features has been counterproductive. Instead, it is suggested, the number of significant features can be in the thousands…

统计理论 · 数学 2014-07-10 Peter Hall , Jiashun Jin , Hugh Miller

The beneficial effects of treatments vary across individuals in most studies. Treatment heterogeneity motivates practitioners to search for the optimal policy based on personal characteristics. A long-standing common practice in policy…

统计理论 · 数学 2025-01-06 Xuqiao Li , Ying Yan

An important first step in computational SAR modeling is to transform the compounds into a representation that can be processed by predictive modeling techniques. This is typically a feature vector where each feature indicates the presence…

计算工程、金融与科学 · 计算机科学 2015-01-14 Albrecht Zimmermann , Björn Bringmann , Luc De Raedt

Feature selection, as a data preprocessing strategy, has been proven to be effective and efficient in preparing data (especially high-dimensional data) for various data mining and machine learning problems. The objectives of feature…

机器学习 · 计算机科学 2018-08-28 Jundong Li , Kewei Cheng , Suhang Wang , Fred Morstatter , Robert P. Trevino , Jiliang Tang , Huan Liu

This paper investigates the effects of data size and frequency range on distributional semantic models. We compare the performance of a number of representative models for several test settings over data of varying sizes, and over test…

计算与语言 · 计算机科学 2016-09-28 Magnus Sahlgren , Alessandro Lenci

The selection of the assumed effect size (AES) critically determines the duration of an experiment, and hence its accuracy and efficiency. Traditionally, experimenters determine AES based on domain knowledge. However, this method becomes…

机器学习 · 计算机科学 2025-04-15 Yu Liu , Runzhe Wan , James McQueen , Doug Hains , Jinxiang Gu , Rui Song

Excluding irrelevant features in a pattern recognition task plays an important role in maintaining a simpler machine learning model and optimizing the computational efficiency. Nowadays with the rise of large scale datasets, feature…

机器学习 · 计算机科学 2018-04-17 Saman Sadeghyan

This research paper investigates the effectiveness of simple linear models versus complex machine learning techniques in breast cancer diagnosis, emphasizing the importance of interpretability and computational efficiency in the medical…

机器学习 · 计算机科学 2023-06-06 Muhammad Arbab Arshad , Sakib Shahriar , Khizar Anjum

Machine learning is capable of discriminating phases of matter, and finding associated phase transitions, directly from large data sets of raw state configurations. In the context of condensed matter physics, most progress in the field of…

统计力学 · 物理学 2017-12-06 Pedro Ponte , Roger G. Melko

This paper presents a method based on a kernel dictionary learning algorithm for segmenting brain tumor regions in magnetic resonance images (MRI). A set of first-order and second-order statistical feature vectors are extracted from patches…

计算机视觉与模式识别 · 计算机科学 2023-10-18 Seyedeh Mahya Mousavi , Mohammad Mostafavi