English
Related papers

Related papers: Weighted Diversified Sampling for Efficient Data-D…

200 papers

We develop a model-based methodology for integrating gene-set information with an experimentally-derived gene list. The methodology uses a previously reported sampling model, but takes advantage of natural constraints in the…

Methodology · Statistics 2015-06-02 Zhishi Wang , Qiuling He , Bret Larget , Michael A. Newton

An applied problem facing all areas of data science is harmonizing data sources. Joining data from multiple origins with unmapped and only partially overlapping features is a prerequisite to developing and testing robust, generalizable…

Leveraging the vast genetic diversity within microbiomes offers unparalleled insights into complex phenotypes, yet the task of accurately predicting and understanding such traits from genomic data remains challenging. We propose a framework…

Genomics · Quantitative Biology 2025-03-05 Zhufeng Li , Sandeep S Cranganore , Nicholas Youngblut , Niki Kilbertus

Most distributed sensing methods assume that the expected value of sensed information is same for all agents ignoring differences in sensor capabilities due to, for example, environmental factors and sensors quality and condition. In this…

Optimization and Control · Mathematics 2015-02-19 John Daniel Peterson , Tansel Yucelen , Girish Chowdhary , Suresh Kannan

Rich phenomena from complex systems have long intrigued researchers, and yet modeling system micro-dynamics and inferring the forms of interaction remain challenging for conventional data-driven approaches, being generally established by…

Statistical Mechanics · Physics 2020-11-13 Seungwoong Ha , Hawoong Jeong

We consider a method to jointly estimate sparse precision matrices and their underlying graph structures using dependent high-dimensional datasets. We present a penalized maximum likelihood estimator which encourages both sparsity and…

Applications · Statistics 2016-08-22 Adria Caballe , Natalia Bochkina , Claus Mayer

Models obtained by decision tree induction techniques excel in being interpretable.However, they can be prone to overfitting, which results in a low predictive performance. Ensemble techniques are able to achieve a higher accuracy. However,…

Machine Learning · Statistics 2016-11-18 Gilles Vandewiele , Olivier Janssens , Femke Ongenae , Filip De Turck , Sofie Van Hoecke

Genetic mutations can cause disease by disrupting normal gene function. Identifying the disease-causing mutations from millions of genetic variants within an individual patient is a challenging problem. Computational methods which can…

Machine Learning · Computer Science 2021-06-28 Jun Cheng , Carolin Lawrence , Mathias Niepert

Computing the exact likelihood of data in large Bayesian networks consisting of thousands of vertices is often a difficult task. When these models contain many deterministic conditional probability tables and when the observed values are…

Computation · Statistics 2012-06-26 Ydo Wexler , Dan Geiger

Deep generative models are proficient in generating realistic data but struggle with producing rare samples in low density regions due to their scarcity of training datasets and the mode collapse problem. While recent methods aim to improve…

Computer Vision and Pattern Recognition · Computer Science 2025-01-08 Subeen Lee , Jiyeon Han , Soyeon Kim , Jaesik Choi

Data complexity analysis quantifies the hardness of constructing a predictive model on a given dataset. However, the effectiveness of existing data complexity measures can be challenged by the existence of irrelevant features and feature…

Computational Engineering, Finance, and Science · Computer Science 2023-08-15 Zhendong Sha , Li Zhu , Zijun Jiang , Yuanzhu Chen , Ting Hu

The sharp increase in data-related expenses has motivated research into condensing datasets while retaining the most informative features. Dataset distillation has thus recently come to the fore. This paradigm generates synthetic datasets…

Machine Learning · Computer Science 2024-11-20 Jiawei Du , Xin Zhang , Juncheng Hu , Wenxin Huang , Joey Tianyi Zhou

In the information overloaded web, personalized recommender systems are essential tools to help users find most relevant information. The most heavily-used recommendation frameworks assume user interactions that are characterized by a…

Information Retrieval · Computer Science 2017-03-06 Fatemeh Vahedian , Robin Burke , Bamshad Mobasher

A large number of complex systems find a natural abstraction in the form of weighted networks whose nodes represent the elements of the system and the weighted edges identify the presence of an interaction and its relative strength. In…

Physics and Society · Physics 2009-04-23 M. Angeles Serrano , Marian Boguna , Alessandro Vespignani

Integrative learning of multiple datasets has the potential to mitigate the challenge of small $n$ and large $p$ that is often encountered in analysis of big biomedical data such as genomics data. Detection of weak yet important signals can…

Methodology · Statistics 2022-07-04 Changgee Chang , Zongyu Dai , Jihwan Oh , Qi Long

Exoplanet atmospheric modeling is advancing from chemically diverse one-dimensional (1D) models to three-dimensional (3D) global circulation models (GCMs), which are crucial for interpreting observations from facilities like the James Webb…

Earth and Planetary Astrophysics · Physics 2024-12-06 A. Lira-Barria , J. N. Harvey , T. Konings , R. Baeyens , C. Henríquez , L. Decin , O. Venot , R. Veillet

DNA sequence classification requires not only high predictive accuracy but also the ability to uncover latent site interactions, combinatorial regulation, and epistasis-like higher-order dependencies. Although the standard Transformer…

Machine Learning · Computer Science 2026-03-30 Zhixuan Cao , Yishu Xu , Xuang WU

Single-cell gene expression data are often characterized by large matrices, where the number of cells may be lower than the number of genes of interest. Factorization models have emerged as powerful tools to condense the available…

Methodology · Statistics 2023-05-22 Antonio Canale , Luisa Galtarossa , Davide Risso , Lorenzo Schiavon , Giovanni Toto

Soft sensing infers hard-to-measure data through a large number of easily obtainable variables. However, in complex industrial scenarios, the issue of insufficient data volume persists, which diminishes the reliability of soft sensing.…

Machine Learning · Computer Science 2025-12-23 Zesen Wang , Yonggang Li , Lijuan Lan

We propose a probabilistic model for interpreting gene expression levels that are observed through single-cell RNA sequencing. In the model, each cell has a low-dimensional latent representation. Additional latent variables account for…

Machine Learning · Computer Science 2017-10-18 Romain Lopez , Jeffrey Regier , Michael Cole , Michael Jordan , Nir Yosef