English
Related papers

Related papers: Preventing dataset shift from breaking machine-lea…

200 papers

Investigation of machine learning algorithms robust to changes between the training and test distributions is an active area of research. In this paper we explore a special type of dataset shift which we call class-dependent domain shift.…

Machine Learning · Computer Science 2020-07-13 Tigran Galstyan , Hrant Khachatrian , Greg Ver Steeg , Aram Galstyan

Machine learning is a modern approach to problem-solving and task automation. In particular, machine learning is concerned with the development and applications of algorithms that can recognize patterns in data and use them for predictive…

Machine learning models for medical image analysis often suffer from poor performance on important subsets of a population that are not identified during training or testing. For example, overall performance of a cancer detection model may…

Machine Learning · Computer Science 2019-11-18 Luke Oakden-Rayner , Jared Dunnmon , Gustavo Carneiro , Christopher Ré

Healthcare datasets present many challenges to both machine learning and statistics as their data are typically heterogeneous, censored, high-dimensional and have missing information. Feature selection is often used to identify the…

Machine Learning · Computer Science 2022-07-06 Annette Spooner , Gelareh Mohammadi , Perminder S. Sachdev , Henry Brodaty , Arcot Sowmya

The discovery of patient-specific imaging markers that are predictive of future disease outcomes can help us better understand individual-level heterogeneity of disease evolution. In fact, deep learning models that can provide data-driven…

Computer Vision and Pattern Recognition · Computer Science 2022-08-05 Amar Kumar , Anjun Hu , Brennan Nichyporuk , Jean-Pierre R. Falet , Douglas L. Arnold , Sotirios Tsaftaris , Tal Arbel

The deep learning approach to detecting malicious software (malware) is promising but has yet to tackle the problem of dataset shift, namely that the joint distribution of examples and their labels associated with the test set is different…

Cryptography and Security · Computer Science 2021-12-15 Deqiang Li , Tian Qiu , Shuo Chen , Qianmu Li , Shouhuai Xu

A key obstacle in automated analytics and meta-learning is the inability to recognize when different datasets contain measurements of the same variable. Because provided attribute labels are often uninformative in practice, this task may be…

Machine Learning · Computer Science 2019-09-12 Jonas Mueller , Alex Smola

Tipping points occur in many real-world systems, at which the system shifts suddenly from one state to another. The ability to predict the occurrence of tipping points from time series data remains an outstanding challenge and a major…

Machine Learning · Computer Science 2024-12-10 Chengzuo Zhuge , Jiawei Li , Wei Chen

In machine learning, if the training data is an unbiased sample of an underlying distribution, then the learned classification function will make accurate predictions for new samples. However, if the training data is not an unbiased sample,…

Machine Learning · Computer Science 2019-01-15 Wouter M. Kouw , Marco Loog

Scientific discoveries are often made by finding a pattern or object that was not predicted by the known rules of science. Oftentimes, these anomalous events or objects that do not conform to the norms are an indication that the rules of…

Machine Learning · Computer Science 2026-02-17 Elizabeth G. Campolongo , Yuan-Tang Chou , Ekaterina Govorkova , Wahid Bhimji , Wei-Lun Chao , Chris Harris , Shih-Chieh Hsu , Hilmar Lapp , Mark S. Neubauer , Josephine Namayanja , Aneesh Subramanian , Philip Harris , Advaith Anand , David E. Carlyn , Subhankar Ghosh , Christopher Lawrence , Eric Moreno , Ryan Raikman , Jiaman Wu , Ziheng Zhang , Bayu Adhi , Mohammad Ahmadi Gharehtoragh , Saúl Alonso Monsalve , Marta Babicz , Furqan Baig , Namrata Banerji , William Bardon , Tyler Barna , Tanya Berger-Wolf , Adji Bousso Dieng , Micah Brachman , Quentin Buat , David C. Y. Hui , Phuong Cao , Franco Cerino , Yi-Chun Chang , Shivaji Chaulagain , An-Kai Chen , Deming Chen , Eric Chen , Chia-Jui Chou , Zih-Chen Ciou , Miles Cochran-Branson , Artur Cordeiro Oudot Choi , Michael Coughlin , Matteo Cremonesi , Maria Dadarlat , Peter Darch , Malina Desai , Daniel Diaz , Steven Dillmann , Javier Duarte , Isla Duporge , Urbas Ekka , Saba Entezari Heravi , Hao Fang , Rian Flynn , Geoffrey Fox , Emily Freed , Hang Gao , Jing Gao , Julia Gonski , Matthew Graham , Abolfazl Hashemi , Scott Hauck , James Hazelden , Joshua Henry Peterson , Duc Hoang , Wei Hu , Mirco Huennefeld , David Hyde , Vandana Janeja , Nattapon Jaroenchai , Haoyi Jia , Yunfan Kang , Maksim Kholiavchenko , Elham E. Khoda , Sangin Kim , Aditya Kumar , Bo-Cheng Lai , Trung Le , Chi-Wei Lee , JangHyeon Lee , Shaocheng Lee , Suzan van der Lee , Charles Lewis , Haitong Li , Haoyang Li , Henry Liao , Mia Liu , Xiaolin Liu , Xiulong Liu , Vladimir Loncar , Fangzheng Lyu , Ilya Makarov , Abhishikth Mallampalli , Chen-Yu Mao , Alexander Michels , Alexander Migala , Farouk Mokhtar , Mathieu Morlighem , Min Namgung , Andrzej Novak , Andrew Novick , Amy Orsborn , Anand Padmanabhan , Jia-Cheng Pan , Sneh Pandya , Zhiyuan Pei , Ana Peixoto , George Percivall , Alex Po Leung , Sanjay Purushotham , Zhiqiang Que , Melissa Quinnan , Arghya Ranjan , Dylan Rankin , Christina Reissel , Benedikt Riedel , Dan Rubenstein , Argyro Sasli , Eli Shlizerman , Arushi Singh , Kim Singh , Eric R. Sokol , Arturo Sorensen , Yu Su , Mitra Taheri , Vaibhav Thakkar , Ann Mariam Thomas , Eric Toberer , Chenghan Tsai , Rebecca Vandewalle , Arjun Verma , Ricco C. Venterea , He Wang , Jianwu Wang , Sam Wang , Shaowen Wang , Gordon Watts , Jason Weitz , Andrew Wildridge , Rebecca Williams , Scott Wolf , Yue Xu , Jianqi Yan , Jai Yu , Yulei Zhang , Haoran Zhao , Ying Zhao , Yibo Zhong

Classification data sets with skewed class proportions are called imbalanced. Class imbalance is a problem since most machine learning classification algorithms are built with an assumption of equal representation of all classes in the…

Machine Learning · Computer Science 2022-12-22 Azal Ahmad Khan

Dataset pruning is the process of removing sub-optimal tuples from a dataset to improve the learning of a machine learning model. In this paper, we compared the performance of different algorithms, first on an unpruned dataset and then on…

Machine Learning · Computer Science 2019-01-31 Arun Thundyill Saseendran , Lovish Setia , Viren Chhabria , Debrup Chakraborty , Aneek Barman Roy

Uncertainty-aware machine learners, such as Bayesian neural networks, output a quantification of uncertainty instead of a point prediction. We provide uncertainty-aware learners with a principled framework to characterize, and identify ways…

Machine Learning · Computer Science 2026-04-01 Sabina J. Sloman , Michele Caprio , Samuel Kaski

The underlying paradigm of big data-driven machine learning reflects the desire of deriving better conclusions from simply analyzing more data, without the necessity of looking at theory and models. Is having simply more data always…

Machine Learning · Computer Science 2018-03-05 Patrick Glauner , Petko Valtchev , Radu State

Imitation learning field requires expert data to train agents in a task. Most often, this learning approach suffers from the absence of available data, which results in techniques being tested on its dataset. Creating datasets is a…

Machine Learning · Computer Science 2024-03-04 Nathan Gavenski , Michael Luck , Odinaldo Rodrigues

Bias in data can have unintended consequences that propagate to the design, development, and deployment of machine learning models. In the financial services sector, this can result in discrimination from certain financial instruments and…

Cryptography and Security · Computer Science 2019-11-12 Reginald Bryant , Celia Cintas , Isaac Wambugu , Andrew Kinai , Komminist Weldemariam

Optimal biomarker combinations for treatment-selection can be derived by minimizing total burden to the population caused by the targeted disease and its treatment. However, when multiple biomarkers are present, including all in the model…

Applications · Statistics 2019-06-07 Sayan Dasgupta , Ying Huang

Quantification is the supervised learning task that consists of training predictors of the class prevalence values of sets of unlabelled data, and is of special interest when the labelled data on which the predictor has been trained and the…

Machine Learning · Computer Science 2023-10-10 Pablo González , Alejandro Moreo , Fabrizio Sebastiani

Benchmark datasets play a central role in the organization of machine learning research. They coordinate researchers around shared research problems and serve as a measure of progress towards shared goals. Despite the foundational role of…

Machine Learning · Computer Science 2021-12-06 Bernard Koch , Emily Denton , Alex Hanna , Jacob G. Foster

Data-centric AI is at the center of a fundamental shift in software engineering where machine learning becomes the new software, powered by big data and computing infrastructure. Here software engineering needs to be re-thought where data…

Machine Learning · Computer Science 2022-12-27 Steven Euijong Whang , Yuji Roh , Hwanjun Song , Jae-Gil Lee