中文
相关论文

相关论文: Measuring Quality of DNA Sequence Data via Degrada…

200 篇论文

Classifying samples in incomplete datasets is a common aim for machine learning practitioners, but is non-trivial. Missing data is found in most real-world datasets and these missing values are typically imputed using established methods,…

Rapid advances in human genomics are enabling researchers to gain a better understanding of the role of the genome in our health and well-being, stimulating hope for more effective and cost efficient healthcare. However, this also prompts a…

密码学与安全 · 计算机科学 2018-08-20 Alexandros Mittos , Bradley Malin , Emiliano De Cristofaro

Traditional data quality control methods are based on users experience or previously established business rules, and this limits performance in addition to being a very time consuming process with lower than desirable accuracy. Utilizing…

人工智能 · 计算机科学 2018-10-17 Wei Dai , Kenji Yoshigoe , William Parsley

Effective data processing depends on the quality of the underlying data. However, quality issues such as inconsistencies and uncertainties, can significantly impede the processing and subsequent use of data. Despite the centrality of data…

数据库 · 计算机科学 2026-02-26 Markus Matoni , Arno Kesper , Gabriele Taentzer

The cost of DNA sequencing has resulted in a surge of genetic data being utilised to improve scientific research, clinical procedures, and healthcare delivery in recent years. Since the human genome can uniquely identify an individual, this…

密码学与安全 · 计算机科学 2022-02-11 Sara Jafarbeiki , Raj Gaire , Amin Sakzad , Shabnam Kasra Kermanshahi , Ron Steinfeld

Data is one of the most important assets of the information age, and its societal impact is undisputed. Yet, rigorous methods of assessing the quality of data are lacking. In this paper, we propose a formal definition for the quality of a…

机器学习 · 计算机科学 2020-05-13 Netanel Raviv , Siddharth Jain , Jehoshua Bruck

A mathematical model of genome degradation is proposed that takes into account a variable rate of mutation and increasing number of cells in a developing human organism. The model explains known properties of cancer development, in…

生物物理 · 物理学 2007-05-23 V. N. Binhi

Data-oriented applications, their users, and even the law require data of high quality. Research has divided the rather vague notion of data quality into various dimensions, such as accuracy, consistency, and reputation. To achieve the goal…

数据库 · 计算机科学 2024-12-09 Sedir Mohammed , Lisa Ehrlinger , Hazar Harmouch , Felix Naumann , Divesh Srivastava

This commentary discusses a recently proposed measure of heterogeneity of DNA sequences and compares with the measures of complexity.

adap-org · 物理学 2012-05-07 Wentian Li

Nowadays, people strive to improve the accuracy of deep learning models. However, very little work has focused on the quality of data sets. In fact, data quality determines model quality. Therefore, it is important for us to make research…

机器学习 · 计算机科学 2019-07-01 Tianxing He , Shengcheng Yu , Ziyuan Wang , Jieqiong Li , Zhenyu Chen

Data quality issues have attracted widespread attention due to the negative impacts of dirty data on data mining and machine learning results. The relationship between data quality and the accuracy of results could be applied on the…

数据库 · 计算机科学 2021-04-27 Zhixin Qi , Hongzhi Wang , Jianzhong Li , Hong Gao

We consider a problem of data integration. Consider determining which genes affect a disease. The genes, which we call predictor objects, can be measured in different experiments on the same individual. We address the question of finding…

机器学习 · 统计学 2016-10-04 Xin Gao , Raymond J. Carroll

This study proposes a data condensation method for multivariate kernel density estimation by genetic algorithm. First, our proposed algorithm generates multiple subsamples of a given size with replacement from the original sample. The…

统计方法学 · 统计学 2022-03-04 Kiheiji Nishida

Data warehousing is continuously gaining importance as organizations are realizing the benefits of decision oriented data bases. However, the stumbling block to this rapid development is data quality issues at various stages of data…

数据库 · 计算机科学 2013-10-09 Vinay Kumar , Reema Thareja

Quantization is widely applied in machine learning to reduce computational and storage costs for both data and models. Considering that classification tasks are fundamental to the field, it is crucial to investigate how quantization impacts…

机器学习 · 计算机科学 2025-07-14 Weizhi Lu , Mingrui Chen , Weiyu Li

The gradual patterns that model the complex co-variations of attributes of the form "The more/less X, The more/less Y" play a crucial role in many real world applications where the amount of numerical data to manage is important, this is…

机器学习 · 计算机科学 2020-05-25 Michaël Chirmeni Boujike , Jerry Lonlac , Norbert Tsopze , Engelbert Mephu Nguifo

As sequencing technologies become more affordable and genomic databases expand continuously, the reuse of publicly available sequencing data emerges as a powerful strategy for studying microbial pathogens. Indeed, raw sequencing reads…

定量方法 · 定量生物学 2025-05-16 Damien Richard , Nils Poulicard

Quality of microarray gene expression data has emerged as a new research topic. As in other areas, microarray quality is assessed by comparing suitable numerical summaries across microarrays, so that outliers and trends can be visualized,…

统计方法学 · 统计学 2011-11-10 Julia Brettschneider , Francois Collin , Benjamin M. Bolstad , Terence P. Speed

Sequential data is everywhere, and it can serve as a basis for research that will lead to improved processes. For example, road infrastructure can be improved by identifying bottlenecks in GPS data, or early diagnosis can be improved by…

密码学与安全 · 计算机科学 2020-02-25 Sigal Shaked , Lior Rokach

The use of learning-based techniques to achieve automated software vulnerability detection has been of longstanding interest within the software security domain. These data-driven solutions are enabled by large software vulnerability…

软件工程 · 计算机科学 2023-01-16 Roland Croft , M. Ali Babar , Mehdi Kholoosi