中文
相关论文

相关论文: Measuring Quality of DNA Sequence Data via Degrada…

200 篇论文

Training deep neural models in the presence of corrupted supervision is challenging as the corrupted data points may significantly impact the generalization performance. To alleviate this problem, we present an efficient robust algorithm…

机器学习 · 计算机科学 2021-02-16 Boyang Liu , Mengying Sun , Ding Wang , Pang-Ning Tan , Jiayu Zhou

Machine Learning (ML) models are being increasingly employed for credit risk evaluation, with their effectiveness largely hinging on the quality of the input data. In this paper we investigate the impact of several data quality issues,…

机器学习 · 计算机科学 2025-11-18 Andrea Maurino

In this paper, we explore how modifying data to preserve privacy affects the quality of the patterns discoverable in the data. For any analysis of modified data to be worth doing, the data must be as close to the original as possible.…

人工智能 · 计算机科学 2015-12-25 Sam Fletcher , Md Zahidul Islam

Data missingness and quality are common problems in machine learning, especially for high-stakes applications such as healthcare. Developers often train machine learning models on carefully curated datasets using only high quality data;…

Genomes may be analyzed from an information viewpoint as very long strings, containing functional elements of variable length, which have been assembled by evolution. In this work an innovative information theory based algorithm is…

基因组学 · 定量生物学 2020-09-23 Vincenzo Bonnici , Giuditta Franco , Vincenzo Manca

Background: The analysis of DNA methylation is a key component in the development of personalized treatment approaches. A common way to measure DNA methylation is the calculation of beta values, which are bounded variables of the form M =…

统计方法学 · 统计学 2016-07-26 Leonie Weinhold , Simone Wahl , Matthias Schmid

DNA has immense potential as an emerging data storage medium. The principle of DNA storage is the conversion and flow of digital information between binary code stream, quaternary base, and actual DNA fragments. This process will inevitably…

信息检索 · 计算机科学 2022-10-21 Yun Qin , Fei Zhu , Bo Xi

Recent large language models have been trained on vast datasets, but also often on repeated data, either intentionally for the purpose of upweighting higher quality data, or unintentionally because data deduplication is not perfect and the…

In recent years, several machine learning approaches have been proposed to predict gene expression and epigenetic signals from the DNA sequence alone. These models are often used to deduce, and, to some extent, assess putative new…

基因组学 · 定量生物学 2023-04-26 Laurent Bréhélin

While conventional wisdom suggests that more aggressively filtering data from low-quality sources like Common Crawl always monotonically improves the quality of training data, we find that aggressive filtering can in fact lead to a decrease…

计算与语言 · 计算机科学 2021-10-08 Leo Gao

We consider the problem of detecting and estimating the strength of association between a trait of interest and alleles or haplotypes in a small genomic region (e.g. a gene or a gene complex), when no direct information on that region is…

应用统计 · 统计学 2008-04-11 Rodrigo Labouriau , Poul Sørensen , Helle R. Juul-Madsen

Understanding degradation is crucial for ensuring the longevity and performance of materials, systems, and organisms. To illustrate the similarities across applications, this article provides a review of data-based method in materials…

系统与控制 · 电气工程与系统科学 2025-09-24 Anna Jarosz-Kozyro , Jerzy Baranowski

Genomic foundation models trained on DNA sequences have demonstrated remarkable capabilities across diverse biological tasks, from variant effect prediction to genome design. These models are typically trained on massive, publicly sourced…

基因组学 · 定量生物学 2026-03-31 Charalampos Koilakos , Ioannis Mouratidis , Ilias Georgakopoulos-Soares

For most diseases, building large databases of labeled genetic data is an expensive and time-demanding task. To address this, we introduce genetic Generative Adversarial Networks (gGAN), a semi-supervised approach based on an innovative GAN…

机器学习 · 计算机科学 2020-07-03 Caio Davi , Ulisses Braga-Neto

Objective audio quality measurement systems often use perceptual models to predict the subjective quality scores of processed signals, as reported in listening tests. Most systems map different metrics of perceived degradation into a single…

音频与语音处理 · 电气工程与系统科学 2022-12-12 Pablo M. Delgado , Jürgen Herre

Research on image quality assessment (IQA) remains limited mainly due to our incomplete knowledge about human visual perception. Existing IQA algorithms have been designed or trained with insufficient subjective data with a small degree of…

计算机视觉与模式识别 · 计算机科学 2020-09-29 Lucie Lévêque , Ji Yang , Xiaohan Yang , Pengfei Guo , Kenneth Dasalla , Leida Li , Yingying Wu , Hantao Liu

The appeal of metric evaluation of research impact has attracted considerable interest in recent times. Although the public at large and administrative bodies are much interested in the idea, scientists and other researchers are much more…

数字图书馆 · 计算机科学 2018-04-10 Fionn Murtagh , Michael Orlov , Boris Mirkin

Central to the efficacy of prognostics and health management methods is the acquisition and analysis of degradation data, which encapsulates the evolving health condition of engineering systems over time. Degradation data serves as a rich…

数据库 · 计算机科学 2026-02-06 Fabian Mauthe , Christopher Braun , Julian Raible , Peter Zeiler , Marco F. Huber

We present the GeneScore, a concept of feature reduction for Machine Learning analysis of biomedical data. Using expert knowledge, the GeneScore integrates different molecular data types into a single score. We show that the GeneScore is…

基因组学 · 定量生物学 2021-01-15 Alexander Denker , Anastasia Steshina , Theresa Grooss , Frank Ueckert , Sylvia Nürnberg

Attribute reduction is one of the most important research topics in the theory of rough sets, and many rough sets-based attribute reduction methods have thus been presented. However, most of them are specifically designed for dealing with…

人工智能 · 计算机科学 2021-01-26 Can Gao , Jie Zhoua , Duoqian Miao , Xiaodong Yue , Jun Wan