English
Related papers

Related papers: Boosted top tagging and its interpretation using S…

200 papers

Topological materials discovery has emerged as an important frontier in condensed matter physics. While theoretical classification frameworks have been used to identify thousands of candidate topological materials, experimental…

Substantial progress in spoofing and deepfake detection has been made in recent years. Nonetheless, the community has yet to make notable inroads in providing an explanation for how a classifier produces its output. The dominance of black…

Audio and Speech Processing · Electrical Eng. & Systems 2024-04-29 Wanying Ge , Jose Patino , Massimiliano Todisco , Nicholas Evans

Keylogger detection involves monitoring for unusual system behaviors such as delays between typing and character display, analyzing network traffic patterns for data exfiltration. In this study, we provide a comprehensive analysis for…

Machine Learning · Computer Science 2025-05-23 Monirul Islam Mahmud

A significant degree of misclassification of variable stars through the application of machine learning methods to survey data motivates a search for more reliable and accurate machine learning procedures, especially in light of the very…

Solar and Stellar Astrophysics · Physics 2019-06-18 Refilwe Kgoadi , Chris Engelbrecht , Ian Whittingham , Andrew Tkachenko

Tree boosting is a highly effective and widely used machine learning method. In this paper, we describe a scalable end-to-end tree boosting system called XGBoost, which is used widely by data scientists to achieve state-of-the-art results…

Machine Learning · Computer Science 2016-06-14 Tianqi Chen , Carlos Guestrin

A popular explainable AI (XAI) approach to quantify feature importance of a given model is via Shapley values. These Shapley values arose in cooperative games, and hence a critical ingredient to compute these in an XAI context is a…

Machine Learning · Computer Science 2022-02-25 Chih-Kuan Yeh , Kuan-Yun Lee , Frederick Liu , Pradeep Ravikumar

A significant challenge in the tagging of boosted objects via machine-learning technology is the prohibitive computational cost associated with training sophisticated models. Nevertheless, the universality of QCD suggests that a large…

High Energy Physics - Phenomenology · Physics 2022-07-13 Frédéric A. Dreyer , Radosław Grabarczyk , Pier Francesco Monni

We present a robust deep incremental learning framework for regression tasks on financial temporal tabular datasets which is built upon the incremental use of commonly available tabular and time series prediction models to adapt to…

Machine Learning · Computer Science 2023-10-11 Thomas Wong , Mauricio Barahona

Explaining the predictions of opaque machine learning algorithms is an important and challenging task, especially as complex models are increasingly used to assist in high-stakes decisions such as those arising in healthcare and finance.…

Machine Learning · Computer Science 2022-06-29 David S. Watson

In top quark production, the polarization of top quarks, decided by the chiral structure of couplings, is likely to be modified in the presence of any new physics contribution to the production. Hence the same is a good discriminator for…

High Energy Physics - Phenomenology · Physics 2019-09-25 Rohini Godbole , Monoranjan Guchait , Charanjit K. Khosa , Jayita Lahiri , Seema Sharma , Aravind H. Vijay

A key element in solving real-life data science problems is selecting the types of models to use. Tree ensemble models (such as XGBoost) are usually recommended for classification and regression problems with tabular data. However, several…

Machine Learning · Computer Science 2021-11-24 Ravid Shwartz-Ziv , Amitai Armon

High $p_T$ Higgs production at hadron colliders provides a direct probe of the internal structure of the $gg \to H$ loop with the $H \to b\bar{b}$ decay offering the most statistics due to the large branching ratio. Despite the overwhelming…

High Energy Physics - Phenomenology · Physics 2018-11-05 Joshua Lin , Marat Freytsis , Ian Moult , Benjamin Nachman

This paper discusses an application of Shapley values in the causal inference field, specifically on how to select the top confounder variables for coarsened exact matching method in a scalable way. We use a dataset from an observational…

Machine Learning · Statistics 2022-06-08 Jilei Yang , Wentao Su

In this article, we review recent machine learning methods used in challenging particle identification of heavy-boosted particles at high-energy colliders. Our primary focus is on attention-based Transformer networks. We report the…

High Energy Physics - Phenomenology · Physics 2024-11-19 A. Hammad , Mihoko M Nojiri

Effective and controllable data selection is critical for LLM instruction tuning, especially with massive open-source datasets. Existing approaches primarily rely on instance-level quality scores, or diversity metrics based on embedding…

Computation and Language · Computer Science 2026-01-21 Zihan Niu , Wenping Hu , Junmin Chen , Xiyue Wang , Tong Xu , Ruiming Tang

The growth of domain-specific applications of semantic models, boosted by the recent achievements of unsupervised embedding learning algorithms, demands domain-specific evaluation datasets. In many cases, content-based recommenders being a…

Computation and Language · Computer Science 2020-11-24 Pierangelo Lombardo , Alessio Boiardi , Luca Colombo , Angelo Schiavone , Nicolò Tamagnone

Additive models, such as produced by gradient boosting, and full interaction models, such as classification and regression trees (CART), are widely used algorithms that have been investigated largely in isolation. We show that these models…

The elastic-input neuro tagger and hybrid tagger, combined with a neural network and Brill's error-driven learning, have already been proposed for the purpose of constructing a practical tagger using as little training data as possible.…

Computation and Language · Computer Science 2007-05-23 Masaki Murata , Qing Ma , Hitoshi Isahara

The objective in extreme multi-label learning is to train a classifier that can automatically tag a novel data point with the most relevant subset of labels from an extremely large label set. Embedding based approaches make training and…

Machine Learning · Computer Science 2015-07-13 Kush Bhatia , Himanshu Jain , Purushottam Kar , Prateek Jain , Manik Varma

Supervised machine learning often operates on the data-driven paradigm, wherein internal model parameters are autonomously optimized to converge predicted outputs with the ground truth, devoid of explicitly programming rules or a priori…

Machine Learning · Computer Science 2024-12-12 Daniel Geissler , Bo Zhou , Mengxi Liu , Paul Lukowicz