English
Related papers

Related papers: Missingness-Adaptive Factor Identification in High…

200 papers

The problem of completing a large matrix with lots of missing entries has received widespread attention in the last couple of decades. Two popular approaches to the matrix completion problem are based on singular value thresholding and…

Statistics Theory · Mathematics 2022-04-25 Sohom Bhattacharya , Sourav Chatterjee

Factor analysis is a flexible technique for assessment of multivariate dependence and codependence. Besides being an exploratory tool used to reduce the dimensionality of multivariate data, it allows estimation of common factors that often…

Methodology · Statistics 2020-02-19 Kelly C. M. Gonçalves , Afonso C. B. Silva

The two-sample test is a fundamental problem in statistics with a wide range of applications. In the realm of high-dimensional data, nonparametric methods have gained prominence due to their flexibility and minimal distributional…

Methodology · Statistics 2024-12-24 Zexi Cai , Wenbo Fei , Doudou Zhou

Attribute value extraction refers to the task of identifying values of an attribute of interest from product information. Product attribute values are essential in many e-commerce scenarios, such as customer service robots, product ranking,…

Computation and Language · Computer Science 2021-12-17 Li Yang , Qifan Wang , Zac Yu , Anand Kulkarni , Sumit Sanghai , Bin Shu , Jon Elsas , Bhargav Kanagal

Fairness is steadily becoming a crucial requirement of Machine Learning (ML) systems. A particularly important notion is subgroup fairness, i.e., fairness in subgroups of individuals that are defined by more than one attributes. Identifying…

Machine Learning · Computer Science 2024-04-30 Giorgos Giannopoulos , Dimitris Sacharidis , Nikolas Theologitis , Loukas Kavouras , Ioannis Emiris

Modern surveys with large sample sizes and growing mixed-type questionnaires require robust and scalable analysis methods. In this work, we consider recovering a mixed dataframe matrix, obtained by complex survey sampling, with entries…

Methodology · Statistics 2024-02-07 Xiaojun Mao , Hengfang Wang , Zhonglei Wang , Shu Yang

Ongoing efforts that span over decades show a rise of AI methods for accelerating scientific discovery, yet accelerating discovery in mathematics remains a persistent challenge for AI. Specifically, AI methods were not effective in creation…

Artificial Intelligence · Computer Science 2026-01-30 Michael Shalyt , Uri Seligmann , Itay Beit Halachmi , Ofir David , Rotem Elimelech , Ido Kaminer

Nonlinear causal discovery from observational data imposes strict identifiability assumptions on the formulation of structural equations utilized in the data generating process. The evaluation of structure learning methods under assumption…

Machine Learning · Statistics 2024-12-17 Georg Velev , Stefan Lessmann

Latent feature models (LFM)s are widely employed for extracting latent structures of data. While offering high, parameter estimation is difficult with LFMs because of the combinational nature of latent features, and non-identifiability is a…

Machine Learning · Computer Science 2018-09-27 Ryota Suzuki , Shingo Takahashi , Murtuza Petladwala , Shigeru Kohmoto

Feature attribution is a fundamental task in both machine learning and data analysis, which involves determining the contribution of individual features or variables to a model's output. This process helps identify the most important…

Machine Learning · Computer Science 2023-10-26 Jinfeng Zhong , Elsa Negre

Two prominent challenges in explainability research involve 1) the nuanced evaluation of explanations and 2) the modeling of missing information through baseline representations. The existing literature introduces diverse evaluation…

Machine Learning · Computer Science 2024-12-24 Oren Barkan , Yehonatan Elisha , Jonathan Weill , Noam Koenigstein

This paper introduces the first theoretical framework for quantifying the efficiency and performance gain opportunity size of adaptive inference algorithms. We provide new approximate and exact bounds for the achievable efficiency and…

Machine Learning · Computer Science 2024-02-08 Soheil Hor , Ying Qian , Mert Pilanci , Amin Arbabian

A novel framework is introduced to formalize identifiability in well-specified but ill-posed linear regression models. The framework is distribution-free and accommodates highly correlated features that may or may not relate to the…

Statistics Theory · Mathematics 2026-03-05 Gianluca Finocchio , Tatyana Krivobokova

There are several confounding factors that can reduce the accuracy of gait recognition systems. These factors can reduce the distinctiveness, or alter the features used to characterise gait, they include variations in clothing, lighting,…

Computer Vision and Pattern Recognition · Computer Science 2016-10-25 Christoforos C. Charalambous , Anil A. Bharath

In the field of machine learning, it is still a critical issue to identify and supervise the learned representation without manually intervening or intuition assistance to extract useful knowledge or serve for the downstream tasks. In this…

Machine Learning · Computer Science 2025-12-10 Shiqi Liu , Jingxin Liu , Qian Zhao , Xiangyong Cao , Huibin Li , Deyu Meng , Hongying Meng , Sheng Liu

We introduce an algorithm for identifying interpretable subgroups with elevated treatment effects, given an estimate of individual or conditional average treatment effects (CATE). Subgroups are characterized by ``rule sets'' --…

Machine Learning · Statistics 2025-07-15 Albert Chiu

We extend the recently introduced regularization/Bayesian System Identification procedures to the estimation of time-varying systems. Specifically, we consider an online setting, in which new data become available at given time steps. The…

Systems and Control · Computer Science 2016-09-26 Giulia Prando , Diego Romeres , Alessandro Chiuso

Hierarchical factor models, which include the bifactor model as a special case, are useful in social and behavioural sciences for measuring hierarchically structured constructs. Specifying a hierarchical factor model involves imposing…

Methodology · Statistics 2026-01-06 Jiawei Qiao , Yunxiao Chen , Zhiliang Ying

Decision trees are a popular family of models due to their attractive properties such as interpretability and ability to handle heterogeneous data. Concurrently, missing data is a prevalent occurrence that hinders performance of machine…

Machine Learning · Computer Science 2020-07-01 Pasha Khosravi , Antonio Vergari , YooJung Choi , Yitao Liang , Guy Van den Broeck

We present a very fast algorithm for general matrix factorization of a data matrix for use in the statistical analysis of high-dimensional data via latent factors. Such data are prevalent across many application areas and generate an…