English
Related papers

Related papers: Categorical Exploratory Data Analysis: From Multic…

200 papers

If the aphorism "All models are wrong"- George Box, continues to be true in data analysis, particularly when analyzing real-world data, then we should annotate this wisdom with visible and explainable data-driven patterns. Such annotations…

Machine Learning · Statistics 2020-12-07 Sabrina Enriquez , Fushing Hsieh

Systemic and idiosyncratic patterns in pitching mechanics of 24 top starting pitchers in Major League Baseball (MLB) are extracted and discovered from PITCHf/x database. These evolving patterns across different pitchers or seasons are…

Applications · Statistics 2018-01-30 Fushing Hsieh , Kevin Fujii , Tania Roy , Cho-Jui Hsieh , Brenda McCowan

With histograms as its foundation, we develop Categorical Exploratory Data Analysis (CEDA) under the extreme-$K$ sample problem, and illustrate its universal applicability through four 1D categorical datasets. Given a sizable $K$, CEDA's…

Applications · Statistics 2020-07-31 Elizabeth Chou , Catie McVey , Yin-Chen Hsieh , Sabrina Enriquez , Fushing Hsieh

The PITCHf/x database has allowed the statistical analysis of of Major League Baseball (MLB) to flourish since its introduction in late 2006. Using PITCHf/x, pitches have been classified by hand, requiring considerable effort, or using…

Applications · Statistics 2013-04-08 Michael A. Pane , Samuel L. Ventura , Rebecca C. Steorts , A. C. Thomas

Since Estimation of Distribution Algorithms (EDA) were proposed, many attempts have been made to improve EDAs' performance in the context of global optimization. So far, the studies or applications of multivariate probabilistic model based…

Neural and Evolutionary Computing · Computer Science 2011-11-10 Weishan Dong , Tianshi Chen , Peter Tino , Xin Yao

Based on structured data derived from large complex systems, we computationally further develop and refine a major factor selection protocol by accommodating structural dependency and heterogeneity among many features to unravel data's…

Methodology · Statistics 2022-09-07 Hsieh Fushing , Elizabeth Chou , Ting-Li Chen

We propose a novel algorithm for supervised dimensionality reduction named Manifold Partition Discriminant Analysis (MPDA). It aims to find a linear embedding space where the within-class similarity is achieved along the direction that is…

Machine Learning · Computer Science 2020-11-24 Yang Zhou , Shiliang Sun

Linear Discriminant Analysis (LDA) is a fundamental method for classification. Its simple linear structure facilitates interpretation, and it is naturally suited to multi-class settings. LDA is also closely connected to several classical…

Methodology · Statistics 2026-04-09 Xin Bing , Bingqing Li , Marten Wegkamp

Co-clustering targets on grouping the samples (e.g., documents, users) and the features (e.g., words, ratings) simultaneously. It employs the dual relation and the bilateral information between the samples and features. In many realworld…

Machine Learning · Computer Science 2016-11-18 Ping Li , Jiajun Bu , Chun Chen , Zhanying He , Deng Cai

We propose NECA, a deep representation learning method for categorical data. Built upon the foundations of network embedding and deep unsupervised representation learning, NECA deeply embeds the intrinsic relationship among attribute values…

Machine Learning · Computer Science 2022-05-26 Xiaonan Gao , Sen Wu , Wenjun Zhou

Manifold matching works to identify embeddings of multiple disparate data spaces into the same low-dimensional space, where joint inference can be pursued. It is an enabling methodology for fusion and inference from multiple and massive…

Machine Learning · Statistics 2012-09-18 Ming Sun , Carey E. Priebe , Minh Tang

Data set composed of categorical features is very common in big data analysis tasks. Since categorical features are usually with a limited number of qualitative possible values, the nested granular cluster effect is prevalent in the…

Machine Learning · Computer Science 2026-01-26 Shenghong Cai , Yiqun Zhang , Xiaopeng Luo , Yiu-Ming Cheung , Hong Jia , Peng Liu

Under any Multiclass Classification (MCC) setting defined by a collection of labeled point-cloud specified by a feature-set, we extract only stochastic partial orderings from all possible triplets of point-cloud without explicitly measuring…

Machine Learning · Statistics 2021-04-16 Fushing Hsieh , Xiaodong Wang

Matrix factor model has been growing popular in scientific fields such as econometrics, which serves as a two-way dimension reduction tool for matrix sequences. In this article, we for the first time propose the matrix elliptical factor…

Methodology · Statistics 2022-03-29 ZeYu Li , Yong He , Xinbing Kong , Xinsheng Zhang

Interacting, self-propelled particles such as epithelial cells can dynamically self-organize into complex multicellular patterns, which are challenging to classify without a priori information. Classically, different phases and phase…

Quantitative Methods · Quantitative Biology 2021-01-19 Dhananjay Bhaskar , William Y. Zhang , Ian Y. Wong

In a world abundant with diverse data arising from complex acquisition techniques, there is a growing need for new data analysis methods. In this paper we focus on high-dimensional data that are organized into several hierarchical datasets.…

Machine Learning · Computer Science 2021-04-06 Lior Aloni , Omer Bobrowski , Ronen Talmon

Through Alzheimer's Disease Neuroimaging Initiative (ADNI), time-to-event data: from the pre-dementia state of mild cognitive impairment (MCI) to the diagnosis of Alzheimer's disease (AD), is collected and analyzed by explicitly unraveling…

Applications · Statistics 2022-11-30 Shuting Liao , Fushing Hsieh

In this paper, a novel feature selection method is presented, which is based on Class-Separability (CS) strategy and Data Envelopment Analysis (DEA). To better capture the relationship between features and the class, class labels are…

Machine Learning · Computer Science 2015-02-03 Yishi Zhang , Chao Yang , Anrong Yang , Chan Xiong , Xingchi Zhou , Zigang Zhang

Principal component analysis (PCA) defines a reduced space described by PC axes for a given multidimensional-data sequence to capture the variations of the data. In practice, we need multiple data sequences that accurately obey individual…

Methodology · Statistics 2021-04-19 Ikuo Fukuda , Kei Moritsugu

Multisource domain adaptation (MDA) aims to use multiple source datasets with available labels to infer labels on a target dataset without available labels for target supervision. Prior works on MDA in the literature is ad-hoc as the…

Computer Vision and Pattern Recognition · Computer Science 2025-12-23 Alexander M. Glandon , Khan M. Iftekharuddin
‹ Prev 1 2 3 10 Next ›