English
Related papers

Related papers: Highly Imbalanced Regression with Tabular Data in …

200 papers

As one of the central tasks in machine learning, regression finds lots of applications in different fields. An existing common practice for solving regression problems is the mean square error (MSE) minimization approach or its regularized…

Machine Learning · Statistics 2022-11-24 Jirong Yi , Qiaosheng Zhang , Zhen Chen , Qiao Liu , Wei Shao , Yusen He , Yaohua Wang

Imbalanced and small data regimes are pervasive in domains such as rare disease imaging, genomics, and disaster response, where labeled samples are scarce and naive augmentation often introduces artifacts. Existing solutions such as…

Machine Learning · Computer Science 2025-09-17 J. Cha , J. Lee , J. Cho , J. Shin

Class imbalance is a common problem in the case of real-world object detection and classification tasks. Data of some classes is abundant making them an over-represented majority, and data of other classes is scarce, making them an…

Computer Vision and Pattern Recognition · Computer Science 2017-03-24 Salman H. Khan , Munawar Hayat , Mohammed Bennamoun , Ferdous Sohel , Roberto Togneri

When estimating a regression model, we might have data where some labels are missing, or our data might be biased by a selection mechanism. When the response or selection mechanism is ignorable (i.e., independent of the response variable…

Statistics Theory · Mathematics 2023-08-22 Philip Boeken , Noud de Kroon , Mathijs de Jong , Joris M. Mooij , Onno Zoeter

Existing regression models tend to fall short in both accuracy and uncertainty estimation when the label distribution is imbalanced. In this paper, we propose a probabilistic deep learning model, dubbed variational imbalanced regression…

Machine Learning · Computer Science 2024-11-12 Ziyan Wang , Hao Wang

For multiple index models, it has recently been shown that the sliced inverse regression (SIR) is consistent for estimating the sufficient dimension reduction (SDR) space if and only if $\rho=\lim\frac{p}{n}=0$, where $p$ is the dimension…

Statistics Theory · Mathematics 2018-06-19 Qian Lin , Zhigen Zhao , Jun S. Liu

We study the linear ill-posed inverse problem with noisy data in the statistical learning setting. Approximate reconstructions from random noisy data are sought with general regularization schemes in Hilbert scale. We discuss the rates of…

Statistics Theory · Mathematics 2024-04-09 Abhishake Rastogi , Peter Mathé

Weak-lensing peak counts provide a straightforward way to constrain cosmology by linking local maxima of the lensing signal to the mass function. Recent applications to data have already been numerous and fruitful. However, the importance…

Cosmology and Nongalactic Astrophysics · Physics 2018-06-13 Chieh-An Lin , Martin Kilbinger

Implicit Neural representations (INRs) are widely used for scientific data reduction and visualization by modeling the function that maps a spatial location to a data value. Without any prior knowledge about the spatial distribution of…

Graphics · Computer Science 2024-02-22 Haoyu Li , Han-Wei Shen

Deep convolutional neural networks often perform poorly when faced with datasets that suffer from quantity imbalances and classification difficulties. Despite advances in the field, existing two-stage approaches still exhibit dataset bias…

Machine Learning · Computer Science 2023-03-16 Liang Xu , Yi Cheng , Fan Zhang , Bingxuan Wu , Pengfei Shao , Peng Liu , Shuwei Shen , Peng Yao , Ronald X. Xu

Automated analysis of tissue sections allows a better understanding of disease biology and may reveal biomarkers that could guide prognosis or treatment selection. In digital pathology, less abundant cell types can be of biological…

Image and Video Processing · Electrical Eng. & Systems 2021-02-24 Yeman Brhane Hagos , Catherine SY Lecat , Dominic Patel , Lydia Lee , Thien-An Tran , Manuel Rodriguez- Justo , Kwee Yong , Yinyin Yuan

Multiple Imputation (MI) is one of the most popular approaches to addressing missing values in questionnaires and surveys. MI with multivariate imputation by chained equations (MICE) allows flexible imputation of many types of data. In…

Methodology · Statistics 2023-04-24 Edoardo Costantini , Kyle M. Lang , Klaas Sijtsma , Tim Reeskens

Hyperspectral image fusion (HIF) is critical to a wide range of applications in remote sensing and many computer vision applications. Most traditional HIF methods assume that the observation model is predefined or known. However, in real…

Computer Vision and Pattern Recognition · Computer Science 2021-04-01 Wu Wang , Yue Huang , Xinhao Ding

Marginal structural models (MSMs) with inverse probability weighting offer an approach to estimating causal effects of treatment sequences on repeated outcome measures in the presence of time-varying confounding and dependent censoring.…

Methodology · Statistics 2018-07-02 Sean Yiu , Li Su

Class-Incremental Learning (CIL) trains a model to continually recognize new classes from non-stationary data while retaining learned knowledge. A major challenge of CIL arises when applying to real-world data characterized by non-uniform…

Computer Vision and Pattern Recognition · Computer Science 2024-04-02 Jiangpeng He , Fengqing Zhu

Credit scoring is a systematic approach to evaluate a borrower's probability of default (PD) on a bank loan. The data associated with such scenarios are characteristically imbalanced, complicating binary classification owing to the…

Machine Learning · Computer Science 2025-01-22 Xia Li , Hanghang Zheng , Kunpeng Tao , Mao Mao

In this paper, we propose a sequential directional importance sampling (SDIS) method for rare event estimation. SDIS expresses a small failure probability in terms of a sequence of auxiliary failure probabilities, defined by magnifying the…

Computation · Statistics 2022-02-14 Kai Cheng , Iason Papaioannou , Zhenzhou Lu , Xiaobo Zhang , Yanping Wang

Real-world object classes appear in imbalanced ratios. This poses a significant challenge for classifiers which get biased towards frequent classes. We hypothesize that improving the generalization capability of a classifier should improve…

Computer Vision and Pattern Recognition · Computer Science 2019-01-24 Munawar Hayat , Salman Khan , Waqas Zamir , Jianbing Shen , Ling Shao

Based on previous work, we assess the use of NIR-HSI images for calibrating models on two datasets, focusing on protein content regression and grain variety classification. Limited reference data for protein content is expanded by…

Computer Vision and Pattern Recognition · Computer Science 2023-11-08 Ole-Christian Galbo Engstrøm , Erik Schou Dreier , Birthe Møller Jespersen , Kim Steenstrup Pedersen

In this paper we consider the problem of recovering a high dimensional data matrix from a set of incomplete and noisy linear measurements. We introduce a new model that can efficiently restrict the degrees of freedom of the problem and is…

Information Theory · Computer Science 2012-11-22 Mohammad Golbabaee , Pierre Vandergheynst