中文
相关论文

相关论文: Measuring Quality of DNA Sequence Data via Degrada…

200 篇论文

When machine learning models encounter data which is out of the distribution on which they were trained they have a tendency to behave poorly, most prominently over-confidence in erroneous predictions. Such behaviours will have disastrous…

机器学习 · 计算机科学 2021-06-25 Jack Dymond

Genetic data collection has become ubiquitous, producing genetic information about health, ancestry, and social traits. However, unregulated use, especially amid evolving scientific understanding, poses serious privacy and discrimination…

计算机与社会 · 计算机科学 2025-06-03 Vivek Ramanan , Ria Vinod , Cole Williams , Sohini Ramachandran , Suresh Venkatasubramanian

Current metagenomic analysis algorithms require significant computing resources, can report excessive false positives (type I errors), may miss organisms (type II errors / false negatives), or scale poorly on large datasets. This paper…

数据库 · 计算机科学 2015-01-23 Ashley Mae Conard , Stephanie Dodson , Jeremy Kepner , Darrell Ricke

DNA-based storage is an emerging technology that enables digital information to be archived in DNA molecules. This method enjoys major advantages over magnetic and optical storage solutions such as exceptional information density, enhanced…

信息论 · 计算机科学 2024-03-13 Daniella Bar-Lev , Itai Orr , Omer Sabary , Tuvi Etzion , Eitan Yaakobi

Data quality is vital for user experience in products reliant on data. As solutions for data quality problems, researchers have developed various taxonomies for different types of issues. However, although some of the existing taxonomies…

数据库 · 计算机科学 2024-05-28 Qiaolin Qin , Heng Li , Ettore Merlo

Genetic association data from national biobanks and large-scale association studies have provided new prospects for understanding the genetic evolution of complex traits and diseases in humans. In turn, genomes from ancient human…

种群与进化 · 定量生物学 2021-08-06 Evan K. Irving-Pease , Rasa Muktupavela , Michael Dannemann , Fernando Racimo

Large-scale statistical analysis of data sets associated with genome sequences plays an important role in modern biology. A key component of such statistical analyses is the computation of $p$-values and confidence bounds for statistics…

应用统计 · 统计学 2011-01-06 Peter J. Bickel , Nathan Boley , James B. Brown , Haiyan Huang , Nancy R. Zhang

Over the last decades, hand-crafted feature extractors have been used to encode image visual properties into feature vectors. Recently, data-driven feature learning approaches have been successfully explored as alternatives for producing…

计算机视觉与模式识别 · 计算机科学 2020-11-25 Érico M. Pereira , Ricardo da S. Torres , Jefersson A. dos Santos

DNA sequencing is revolutionising the field of medicine. DNA sequencers, the machines which perform DNA sequencing, have evolved from the size of a fridge to that of a mobile phone over the last two decades. The cost of sequencing a human…

基因组学 · 定量生物学 2021-01-14 Hasindu Gamaarachchi

Synthetic data generation is a promising technique to facilitate the use of sensitive data while mitigating the risk of privacy breaches. However, for synthetic data to be useful in downstream analysis tasks, it needs to be of sufficient…

机器学习 · 统计学 2024-08-26 Thom Benjamin Volker , Peter-Paul de Wolf , Erik-Jan van Kesteren

De novo molecule generation can suffer from data inefficiency; requiring large amounts of training data or many sampled data points to conduct objective optimization. The latter is a particular disadvantage when combining deep generative…

计算工程、金融与科学 · 计算机科学 2025-10-30 Morgan Thomas , Noel M. O'Boyle , Andreas Bender , Chris De Graaf

We study the data deletion problem for convex models. By leveraging techniques from convex optimization and reservoir sampling, we give the first data deletion algorithms that are able to handle an arbitrarily long sequence of adversarial…

机器学习 · 统计学 2020-07-07 Seth Neel , Aaron Roth , Saeed Sharifi-Malvajerdi

The availability of genomic data is essential to progress in biomedical research, personalized medicine, etc. However, its extreme sensitivity makes it problematic, if not outright impossible, to publish or share it. As a result, several…

基因组学 · 定量生物学 2022-01-19 Bristena Oprisanu , Georgi Ganev , Emiliano De Cristofaro

DNA sequencing allows for the determination of the genetic code of an organism, and therefore is an indispensable tool that has applications in Medicine, Life Sciences, Evolutionary Biology, Food Sciences and Technology, and Agriculture. In…

量子物理 · 物理学 2023-04-24 Nouhaila Innan , Muhammad Al-Zafar Khan

In this paper, we consider the outer channel for DNA-based data storage. When transmitting over the outer channel, each DNA string is treated as a unit/symbol that would be either correctly received, or erased, or corrupted by uniformly…

信息论 · 计算机科学 2025-09-23 Xuan He , Yi Ding , Kui Cai , Guanghui Song , Bin Dai , Xiaohu Tang

Sequential querying of differentially private mechanisms degrades the overall privacy level. In this paper, we answer the fundamental question of characterizing the level of overall privacy degradation as a function of the number of queries…

数据结构与算法 · 计算机科学 2015-12-08 Peter Kairouz , Sewoong Oh , Pramod Viswanath

Current techniques in sequencing a genome allow a service provider (e.g. a sequencing company) to have full access to the genome information, and thus the privacy of individuals regarding their lifetime secret is violated. In this paper, we…

基因组学 · 定量生物学 2018-11-28 Ali Gholami , Mohammad Ali Maddah-Ali , Seyed Abolfazl Motahari

Classification for degraded images having various levels of degradation is very important in practical applications. This paper proposes a convolutional neural network to classify degraded images by using a restoration network and an…

计算机视觉与模式识别 · 计算机科学 2020-06-16 Kazuki Endo , Masayuki Tanaka , Masatoshi Okutomi

It is well known that data is critical for training neural networks. Lot have been written about quantities of data required to train networks well. However, there is not much publications on how data quality effects convergence of such…

计算机视觉与模式识别 · 计算机科学 2020-02-11 Subrata Goswami

The US Census Bureau will deliberately corrupt data sets derived from the 2020 US Census, enhancing the privacy of respondents while potentially reducing the precision of economic analysis. To investigate whether this trade-off is…

计量经济学 · 经济学 2024-02-13 Anish Agarwal , Rahul Singh