中文
相关论文

相关论文: Fast Convergence on Perfect Classification for Fun…

200 篇论文

We develop a systematic, omnibus approach to goodness-of-fit testing for parametric distributional models when the variable of interest is only partially observed due to censoring and/or truncation. In many such designs, tests based on the…

统计方法学 · 统计学 2026-02-10 Juan Carlos Escanciano , Jacobo de Uña-Álvarez

We extend the problem of obtaining an estimator for the finite population mean parameter incorporating complete auxiliary information through calibration estimation in survey sampling but considering a functional data framework. The…

统计理论 · 数学 2013-02-06 Santiago Gallón , Jean-Michel Loubes , Fabrice Gamboa

Predicting the mean-field Hamiltonian matrix in density functional theory is a fundamental formulation to leverage machine learning for solving molecular science problems. Yet, its applicability is limited by insufficient labeled data for…

机器学习 · 计算机科学 2024-06-06 He Zhang , Chang Liu , Zun Wang , Xinran Wei , Siyuan Liu , Nanning Zheng , Bin Shao , Tie-Yan Liu

We study the matrix completion problem that leverages hierarchical similarity graphs as side information in the context of recommender systems. Under a hierarchical stochastic block model that well respects practically-relevant social…

信息论 · 计算机科学 2021-09-14 Junhyung Ahn , Adel Elmahdy , Soheil Mohajer , Changho Suh

Motivated by results on generic-case complexity in group theory, we apply the ideas of effective Baire category and effective measure theory to study complexity classes of functions which are "fractionally computable" by a partial…

群论 · 数学 2007-06-30 Ilya Kapovich , Paul Schupp

We introduce a kernel-based goodness-of-fit test for censored data, where observations may be missing in random time intervals: a common occurrence in clinical trials and industrial life-testing. The test statistic is straightforward to…

统计方法学 · 统计学 2018-10-11 Tamara Fernández , Arthur Gretton

Machine learning classification tasks often benefit from predicting a set of possible labels with confidence scores to capture uncertainty. However, existing methods struggle with the high-dimensional nature of the data and the lack of…

机器学习 · 计算机科学 2024-07-08 Rui Luo , Zhixin Zhou

The aim of ordinal classification is to predict the ordered labels of the output from a set of observed inputs. Interval-valued data refers to data in the form of intervals. For the first time, interval-valued data and interval-valued…

统计方法学 · 统计学 2023-11-06 Aleix Alcacer , Marina Martínez-Garcia , Irene Epifanio

We examine overlapping clustering schemes with functorial constraints, in the spirit of Carlsson--Memoli. This avoids issues arising from the chaining required by partition-based methods. Our principal result shows that any clustering…

机器学习 · 计算机科学 2016-08-16 Jared Culbertson , Dan P. Guralnik , Jakob Hansen , Peter F. Stiller

Classification of high-dimensional low sample size (HDLSS) data poses a challenge in a variety of real-world situations, such as gene expression studies, cancer research, and medical imaging. This article presents the development and…

机器学习 · 统计学 2026-05-27 Jyotishka Ray Choudhury , Aytijhya Saha , Sarbojit Roy , Subhajit Dutta

We study multivariate integration and approximation for functions belonging to a weighted reproducing kernel Hilbert space based on half-period cosine functions in the worst-case setting. The weights in the norm of the function space depend…

数值分析 · 数学 2015-11-23 Christian Irrgeher , Peter Kritzer , Friedrich Pillichshammer

One of the key challenges of collaborative machine learning, without data sharing, is multimodal data heterogeneity in real-world settings. While Federated Learning (FL) enables model training across multiple clients, existing frameworks,…

机器学习 · 计算机科学 2025-10-16 Alejandro Guerra-Manzanares , Omar El-Herraoui , Michail Maniatakos , Farah E. Shamout

Cross-silo federated learning (FL) enables decentralized organizations to collaboratively train models while preserving data privacy and has made significant progress in medical image classification. One common assumption is task…

机器学习 · 计算机科学 2024-06-28 Zhaobin Sun , Nannan Wu , Junjie Shi , Li Yu , Xin Yang , Kwang-Ting Cheng , Zengqiang Yan

Data is fundamental to the training of language models (LM). Recent research has been dedicated to data efficiency, which aims to maximize performance by selecting a minimal or optimal subset of training data. Techniques such as data…

计算与语言 · 计算机科学 2025-06-30 Yalun Dai , Yangyu Huang , Xin Zhang , Wenshan Wu , Chong Li , Wenhui Lu , Shijie Cao , Li Dong , Scarlett Li

Data heterogeneity across clients in federated learning (FL) settings is a widely acknowledged challenge. In response, personalized federated learning (PFL) emerged as a framework to curate local models for clients' tasks. In PFL, a common…

机器学习 · 计算机科学 2023-12-27 Yutong Dai , Zeyuan Chen , Junnan Li , Shelby Heinecke , Lichao Sun , Ran Xu

Clustering functional data in the presence of phase variation is challenging, as temporal misalignment can obscure intrinsic shape differences and degrade clustering performance. Most existing approaches treat registration and clustering as…

机器学习 · 统计学 2026-04-30 Xinyang Xiong , Siyuan jiang , Pengcheng Zeng

As the development of measuring instruments and computers has accelerated the collection of massive amounts of data, functional data analysis (FDA) has experienced a surge of attention. The FDA methodology treats longitudinal data as a set…

统计方法学 · 统计学 2024-07-09 Tomoya Wakayama , Hidetoshi Matsui

The problem of supervised classification (or discrimination) with functional data is considered, with a special interest on the popular k-nearest neighbors (k-NN) classifier. First, relying on a recent result by Cerou and Guyader (2006), we…

机器学习 · 统计学 2008-06-18 Amparo Baillo , Antonio Cuevas

Density-corrected density functional theory (DC-DFT) is enjoying substantial success in improving semilocal DFT calculations in a wide variety of chemical problems. This paper provides the formal theoretical framework and assumptions for…

化学物理 · 物理学 2019-08-19 Stefan Vuckovic , Suhwan Song , John Kozlowski , Eunji Sim , Kieron Burke

Real-world object classes appear in imbalanced ratios. This poses a significant challenge for classifiers which get biased towards frequent classes. We hypothesize that improving the generalization capability of a classifier should improve…

计算机视觉与模式识别 · 计算机科学 2019-01-24 Munawar Hayat , Salman Khan , Waqas Zamir , Jianbing Shen , Ling Shao