English
Related papers

Related papers: Pooling Design and Bias Correction in DNA Library …

200 papers

Efficient and accurate product relevance assessment is critical for user experiences and business success. Training a proficient relevance assessment model requires high-quality query-product pairs, often obtained through negative sampling…

Information Retrieval · Computer Science 2024-08-20 Xiaochen Wang , Xiao Xiao , Ruhan Zhang , Xuan Zhang , Taesik Na , Tejaswi Tenneti , Haixun Wang , Fenglong Ma

Model-based clustering is a powerful tool that is often used to discover hidden structure in data by grouping observational units that exhibit similar response values. Recently, clustering methods have been developed that permit…

Methodology · Statistics 2025-06-24 Sally Paganin , Garritt L. Page , Fernando Andrés Quintana

This version is ***superseded*** by a full version that can be found at http://www.itu.dk/people/pagh/papers/mining-jour.pdf, which contains stronger theoretical results and fixes a mistake in the reporting of experiments. Abstract:…

Data Structures and Algorithms · Computer Science 2010-02-17 Andrea Campagna , Rasmus Pagh

Pairwise learning strategies are prevalent for optimizing recommendation models on implicit feedback data, which usually learns user preference by discriminating between positive (i.e., clicked by a user) and negative items (i.e., obtained…

Information Retrieval · Computer Science 2023-02-17 Xiao Chen , Wenqi Fan , Jingfan Chen , Haochen Liu , Zitao Liu , Zhaoxiang Zhang , Qing Li

Pooling designs are standard experimental tools in many biotechnical applications. It is well-known that all famous pooling designs are constructed from mathematical structures by the "containment matrix" method. In particular, Macula's…

Combinatorics · Mathematics 2011-05-16 Jun Guo , Kaishun Wang

In this paper, fundamental limits in sequencing of a set of closely related DNA molecules are addressed. This problem is called pooled-DNA sequencing which encompasses many interesting problems such as haplotype phasing, metageomics, and…

Information Theory · Computer Science 2016-04-20 Amir Najafi , Damoun Nashta-ali , Seyed Abolfazl Motahari , Mehrdad Khani , Babak H. Khalaj , Hamid R. Rabiee

Accurate uncertainty estimates are essential for deploying deep object detectors in safety-critical systems. The development and evaluation of probabilistic object detectors have been hindered by shortcomings in existing performance…

Computer Vision and Pattern Recognition · Computer Science 2022-12-09 Georg Hess , Christoffer Petersson , Lennart Svensson

Recommender systems often struggle with long-tail distributions and limited item catalog exposure, where a small subset of popular items dominates recommendations. This challenge is especially critical in large-scale online retail settings…

Information Retrieval · Computer Science 2026-02-10 Vasileios Karlis , Ezgi Yıldırım , David Vos , Maarten de Rijke

Recent studies have shown that recommendation systems commonly suffer from popularity bias. Popularity bias refers to the problem that popular items (i.e., frequently rated items) are recommended frequently while less popular items are…

Information Retrieval · Computer Science 2022-03-01 Mohammadmehdi Naghiaei , Hossein A. Rahmani , Mahdi Dehghan

We address the issue of binary classification from positive and unlabeled data (PU classification) with a selection bias in the positive data. During the observation process, (i) a sample is exposed to a user, (ii) the user then returns the…

Machine Learning · Computer Science 2023-03-09 Masahiro Kato , Shuting Wu , Kodai Kureishi , Shota Yasui

Domain-specific image collections present potential value in various areas of science and business but are often not curated nor have any way to readily extract relevant content. To employ contemporary supervised image analysis methods on…

Machine Learning · Computer Science 2020-03-10 Sara Mousavi , Dylan Lee , Tatianna Griffin , Dawnie Steadman , Audris Mockus

Classifiers are biased when trained on biased datasets. As a remedy, we propose Learning to Split (ls), an algorithm for automatic bias detection. Given a dataset with input-label pairs, ls learns to split this dataset so that predictors…

Machine Learning · Computer Science 2022-07-22 Yujia Bao , Regina Barzilay

Progressive filtering is a simple way to perform hierarchical classification, inspired by the behavior that most humans put into practice while attempting to categorize an item according to an underlying taxonomy. Each node of the taxonomy…

Artificial Intelligence · Computer Science 2016-11-04 Giuliano Armano

We study the problem of clustering a set of items based on bandit feedback. Each of the $n$ items is characterized by a feature vector, with a possibly large dimension $d$. The items are partitioned into two unknown groups such that items…

Machine Learning · Statistics 2025-03-19 Maximilian Graf , Victor Thuot , Nicolas Verzelen

Standard randomized benchmarking protocols entail sampling from a unitary 2 design, which is not always practical. In this article we examine randomized benchmarking protocols based on subgroups of the Clifford group that are not unitary 2…

Quantum Physics · Physics 2018-06-20 Winton G. Brown , Bryan Eastin

The paper introduces the concept of a cluster structure to define a joint distribution of the sample size and its exchangeable random partitions. The cluster structure allows the probability distribution of the random partitions of a subset…

Methodology · Statistics 2013-10-08 Mingyuan Zhou

Experimental designs with hierarchically-structured errors are pervasive in many biomedical areas; it is important to take into account this hierarchical architecture in order to account for the dispersion and make reliable inferences from…

Applications · Statistics 2023-05-05 Elyas Mouhou , Vincent Audigier , Josselin Noirel

Challenges of assessing complexity and clonality in populations of mixed species arise in diverse areas of modern biology, including estimating diversity and clonality in microbiome populations, measuring patterns of T and B cell clonality,…

Methodology · Statistics 2014-08-07 Yi Liu , Andrew Z. Fire , Scott Boyd , Richard A. Olshen

Bayesian model-based clustering is a widely applied procedure for discovering groups of related observations in a dataset. These approaches use Bayesian mixture models, estimated with MCMC, which provide posterior samples of the model…

Methodology · Statistics 2018-09-24 Ketong Wang , Michael D. Porter

Group testing was conceived during World War II to identify soldiers infected with syphilis using as few tests as possible, and it has attracted renewed interest during the COVID-19 pandemic. A long-standing assumption in the probabilistic…

Social and Information Networks · Computer Science 2022-11-18 Surin Ahn , Wei-Ning Chen , Ayfer Ozgur