English
Related papers

Related papers: YFitter: Maximum likelihood assignment of Y chromo…

200 papers

Motivation: Microarray data has been recently been shown to be efficacious in distinguishing closely related cell types that often appear in the diagnosis of cancer. It is useful to determine the minimum number of genes needed to do such a…

Biological Physics · Physics 2007-05-23 J. M. Deutsch

We propose an algorithm for solving nonlinear convex programs defined in terms of a symmetric positive semidefinite matrix variable $X$. This algorithm rests on the factorization $X=Y Y^T$, where the number of columns of Y fixes the rank of…

Optimization and Control · Mathematics 2010-08-25 M. Journée , F. Bach , P. -A. Absil , R. Sepulchre

DNA samples are often pooled, either by experimental design, or because the sample itself is a mixture. For example, when population allele frequencies are of primary interest, individual samples may be pooled together to lower the cost of…

Quantitative Methods · Quantitative Biology 2013-02-07 Darren Kessner , Tom Turner , John Novembre

Accurate segmentation and classification of nuclei in histology images is critical but challenging due to nuclei heterogeneity, staining variations, and tissue complexity. Existing methods often struggle with limited dataset variability,…

Image and Video Processing · Electrical Eng. & Systems 2025-01-22 Wenhua Zhang , Sen Yang , Meiwei Luo , Chuan He , Yuchen Li , Jun Zhang , Xiyue Wang , Fang Wang

We consider the problem of selecting confounders for adjustment from a potentially large set of covariates, when estimating a causal effect. Recently, the high-dimensional Propensity Score (hdPS) method was developed for this task; hdPS…

Methodology · Statistics 2021-12-17 Asad Haris , Robert Platt

We consider the problem of detecting and estimating the strength of association between a trait of interest and alleles or haplotypes in a small genomic region (e.g. a gene or a gene complex), when no direct information on that region is…

Applications · Statistics 2008-04-11 Rodrigo Labouriau , Poul Sørensen , Helle R. Juul-Madsen

Hierarchical clustering is a fundamental task often used to discover meaningful structures in data, such as phylogenetic trees, taxonomies of concepts, subtypes of cancer, and cascades of particle decays in particle physics. Typically…

Data Structures and Algorithms · Computer Science 2020-10-23 Craig S. Greenberg , Sebastian Macaluso , Nicholas Monath , Ji-Ah Lee , Patrick Flaherty , Kyle Cranmer , Andrew McGregor , Andrew McCallum

Comprehensive discovery of structural variation (SV) in human genomes from DNA sequencing requires the integration of multiple alignment signals including read-pair, split-read and read-depth. However, owing to inherent technical…

Genomics · Quantitative Biology 2014-01-23 Ryan M. Layer , Ira M. Hall , Aaron R. Quinlan

Most successful machine intelligence systems rely on gradient-based learning, which is made possible by backpropagation. Some systems are designed to aid us in interpreting data when explicit goals cannot be provided. These unsupervised…

Machine Learning · Computer Science 2018-06-05 Aditya Ramesh , Yann LeCun

Prior work has shown that Visual Recognition datasets frequently underrepresent bias groups $B$ (\eg Female) within class labels $Y$ (\eg Programmers). This dataset bias can lead to models that learn spurious correlations between class…

Computer Vision and Pattern Recognition · Computer Science 2023-04-28 Maan Qraitem , Kate Saenko , Bryan A. Plummer

The Nystr\"om method is a convenient heuristic method to obtain low-rank approximations to kernel matrices in nearly linear complexity. Existing studies typically use the method to approximate positive semidefinite matrices with low or…

Numerical Analysis · Mathematics 2023-07-13 Jianlin Xia

We develop a highly scalable optimization method called "hierarchical group-thresholding" for solving a multi-task regression model with complex structured sparsity constraints on both input and output spaces. Despite the recent emergence…

Machine Learning · Statistics 2012-08-16 Seunghak Lee , Eric P. Xing

Outcomes of data-driven AI models cannot be assumed to be always correct. To estimate the uncertainty in these outcomes, the uncertainty wrapper framework has been proposed, which considers uncertainties related to model fit, input quality,…

Machine Learning · Computer Science 2022-01-11 Pascal Gerber , Lisa Jöckel , Michael Kläs

Identifying galaxy groups from redshift surveys of galaxies plays an important role in connecting galaxies with the underlying dark matter distribution. Current and future high-$z$ spectroscopic surveys, usually incomplete in redshift…

Astrophysics of Galaxies · Physics 2020-10-14 Kai Wang , H. J. Mo , Cheng Li , Jiacheng Meng , Yangyao Chen

We here introduce a novel classification approach adopted from the nonlinear model identification framework, which jointly addresses the feature selection and classifier design tasks. The classifier is constructed as a polynomial expansion…

Machine Learning · Computer Science 2016-07-29 Aida Brankovic , Alessandro Falsone , Maria Prandini , Luigi Piroddi

Indefinite similarity measures can be frequently found in bio-informatics by means of alignment scores, but are also common in other fields like shape measures in image retrieval. Lacking an underlying vector space, the data are given as…

Machine Learning · Computer Science 2016-04-11 Frank-Michael Schleif , Andrej Gisbrecht , Peter Tino

In this paper a hybrid feature selection method is proposed which takes advantages of wrapper subset evaluation with a lower cost and improves the performance of a group of classifiers. The method uses combination of sample domain filtering…

Machine Learning · Computer Science 2014-03-12 Mehdi Naseriparsa , Amir-Masoud Bidgoli , Touraj Varaee

We consider the problem of setting confidence intervals on a parameter of interest from the maximum-likelihood fit of a physics model to a binned data set with a large number of bins, large event-counts per bin, and in the presence of…

Data Analysis, Statistics and Probability · Physics 2026-02-09 Cristina-Andreea Alexe , Joshua Bendavid , Lorenzo Bianchini , Davide Bruschini

For extremely weak-supervised text classification, pioneer research generates pseudo labels by mining texts similar to the class names from the raw corpus, which may end up with very limited or even no samples for the minority classes.…

Computation and Language · Computer Science 2024-06-18 Letian Peng , Yi Gu , Chengyu Dong , Zihan Wang , Jingbo Shang

B-cell repertoires are characterized by a diverse set of receptors of distinct specificities generated through two processes of somatic diversification: V(D)J recombination and somatic hypermutations. B cell clonal families stem from the…

Populations and Evolution · Quantitative Biology 2024-03-19 Natanael Spisak , Gabriel Athènes , Thomas Dupic , Thierry Mora , Aleksandra M. Walczak