English
Related papers

Related papers: ULV: A robust statistical method for clustered dat…

200 papers

Many complex systems, ranging from migrating cells to animal groups, exhibit stochastic dynamics described by the underdamped Langevin equation. Inferring such an equation of motion from experimental data can provide profound insight into…

Biological Physics · Physics 2026-04-17 David B. Brückner , Pierre Ronceray , Chase P. Broedersz

Automatic sleep stage scoring is crucial for the diagnosis and treatment of sleep disorders. Although deep learning models have advanced the field, many existing models are computationally demanding and designed for single-channel…

Machine Learning · Computer Science 2026-03-02 Zhaowen Wang , Dongdong Zhou , Qi Xu , Fengyu Cong , Mohammad Al-Sa'd , Jenni Raitoharju

Unsupervised feature selection (UFS) is an important task in data engineering. However, most UFS methods construct models from a single perspective and often fail to simultaneously evaluate feature importance and preserve their inherent…

Machine Learning · Computer Science 2025-05-28 Jingjing Liu , Xiansen Ju , Xianchao Xiu , Wanquan Liu

The estimation of covariance matrices of gene expressions has many applications in cancer systems biology. Many gene expression studies, however, are hampered by low sample size and it has therefore become popular to increase sample size by…

The studies of large-scale, high-dimensional data in fields such as genomics and neuroscience have injected new insights into science. Yet, despite advances, they are confronting several challenges, often simultaneously: lack of…

Methodology · Statistics 2024-01-01 Julien Bodelet , Guillaume Blanc , Jiajun Shan , Graciela Muniz Terrera , Oliver Y. Chen

Age-specific probabilities of death provide a snapshot of population mortality at the country level at a given point in time. Due to the high dimensionality of the data, summarising mortality information is essential for various analyses,…

Applications · Statistics 2026-03-30 Pedro Menezes de Araújo , Isobel Claire Gormley , Thomas Brendan Murphy

Advances in data collecting technologies in genomics have significantly increased the need for tools designed to study the genetic basis of many diseases. Effective statistical methods should excel in both prediction accuracy and biomarker…

Methodology · Statistics 2025-11-13 Anthony-Alexander Christidis , Stefan Van Aelst , Ruben Zamar

Clustering analysis is one of the most widely used statistical tools in many emerging areas such as microarray data analysis. For microarray and other high-dimensional data, the presence of many noise variables may mask underlying…

Machine Learning · Statistics 2008-03-26 Benhuai Xie , Wei Pan , Xiaotong Shen

In this article, we propose a class of $L_q$-norm based U-statistics for a family of global testing problems related to high-dimensional data. This includes testing of mean vector and its spatial sign, simultaneous testing of linear model…

Statistics Theory · Mathematics 2023-03-16 Yangfan Zhang , Runmin Wang , Xiaofeng Shao

Accurate classification of computed tomography (CT) images is essential for diagnosis and treatment planning, but existing methods often struggle with the subtle and spatially diverse nature of pathological features. Current approaches…

Image and Video Processing · Electrical Eng. & Systems 2025-07-21 Shravan Venkatraman , Pavan Kumar S , Rakesh Raj Madavan , Chandrakala S

In recent years, functional genomics approaches combining genetic information with bulk RNA-sequencing data have identified the downstream expression effects of disease-associated genetic risk factors through so-called expression…

Over the past decade, the revolution in single-cell sequencing has enabled the simultaneous molecular profiling of various modalities across thousands of individual cells, allowing scientists to investigate the diverse functions of complex…

Computation and Language · Computer Science 2024-12-05 Junhao Liu , Siwei Xu , Lei Zhang , Jing Zhang

Recent studies have raised significant concerns regarding the reliability of current mathematics benchmarks, highlighting issues such as simplistic design and potential data contamination. Consequently, developing a reliable benchmark that…

Computation and Language · Computer Science 2025-08-14 Zijin Hong , Hao Wu , Su Dong , Junnan Dong , Yilin Xiao , Yujing Zhang , Zhu Wang , Feiran Huang , Linyi Li , Hongxia Yang , Xiao Huang

Reasoning is a fundamental capability for solving complex multi-step problems, particularly in visual contexts where sequential step-wise understanding is essential. Existing approaches lack a comprehensive framework for evaluating visual…

Defining and separating cancer subtypes is essential for facilitating personalized therapy modality and prognosis of patients. The definition of subtypes has been constantly recalibrated as a result of our deepened understanding. During…

Machine Learning · Computer Science 2022-07-21 Zheng Chen , Ziwei Yang , Lingwei Zhu , Guang Shi , Kun Yue , Takashi Matsubara , Shigehiko Kanaya , MD Altaf-Ul-Amin

Sequencing costs currently prohibit the application of single-cell mRNA-seq to many biological and clinical analyses. Targeted single-cell mRNA-sequencing reduces sequencing costs by profiling reduced gene sets that capture biological…

Genomics · Quantitative Biology 2022-02-15 Xiaoqiao Chen , Sisi Chen , Matt Thomson

Unsupervised anomaly detection is a promising technique for identifying unusual patterns in data without the need for labeled training examples. This approach is particularly valuable for early case detection in epidemic management,…

Machine Learning · Computer Science 2025-05-06 Ghazal Ghajari , Mithun Kumar PK , Fathi Amsaad

Irregular sampling and high missingness are intrinsic challenges in modeling time series derived from electronic health records (EHRs),where clinical variables are measured at uneven intervals depending on workflow and intervention timing.…

Machine Learning · Computer Science 2025-09-29 Jeong Eul Kwon , Joo Heung Yoon , Hyo Kyung Lee

Disentangled representations have recently been shown to improve fairness, data efficiency and generalisation in simple supervised and reinforcement learning tasks. To extend the benefits of disentangled representations to more complex…

Unsupervised learning on imbalanced data is challenging because, when given imbalanced data, current model is often dominated by the major category and ignores the categories with small amount of data. We develop a latent variable model…

Machine Learning · Computer Science 2016-07-04 Fariba Yousefi , Zhenwen Dai , Carl Henrik Ek , Neil Lawrence