English
Related papers

Related papers: Ask2Me VarHarmonizer: A Python-Based Tool to Harmo…

200 papers

dame-flame is a Python package for performing matching for observational causal inference on datasets containing discrete covariates. This package implements the Dynamic Almost Matching Exactly (DAME) and Fast Large-Scale Almost Matching…

Detecting predictive biomarkers from multi-omics data is important for precision medicine, to improve diagnostics of complex diseases and for better treatments. This needs substantial experimental efforts that are made difficult by the…

Quantitative Methods · Quantitative Biology 2021-06-08 Betül Güvenç Paltun , Samuel Kaski , Hiroshi Mamitsuka

DNA read mapping is a ubiquitous task in bioinformatics, and many tools have been developed to solve the read mapping problem. However, there are two trends that are changing the landscape of readmapping: First, new sequencing technologies…

Genomics · Quantitative Biology 2017-02-09 Jens Quedenfeld , Sven Rahmann

After the completion of human genome sequence was anounced, it is evident that interpretation of DNA sequences is an immediate task to work on. For understanding their signals, improvement of present sequence analysis tools and developing…

Computational Complexity · Computer Science 2007-05-23 Gene Kim , MyungHo Kim

Quantitative organ assessment is an essential step in automated abdominal disease diagnosis and treatment planning. Artificial intelligence (AI) has shown great potential to automatize this process. However, most existing AI algorithms rely…

Conformance checking is a major function of process mining, which allows organizations to identify and alleviate potential deviations from the intended process behavior. To fully leverage its benefits, it is important that conformance…

Software Engineering · Computer Science 2022-09-21 Jana-Rebecca Rehse , Luise Pufahl , Michael Grohs , Lisa-Marie Klein

Objective: To detect and classify features of stigmatizing and biased language in intensive care electronic health records (EHRs) using natural language processing techniques. Materials and Methods: We first created a lexicon and regular…

Computation and Language · Computer Science 2025-07-15 Drew Walker , Annie Thorne , Sudeshna Das , Jennifer Love , Hannah LF Cooper , Melvin Livingston , Abeed Sarker

Electronic health record (EHR) systems are used extensively throughout the healthcare domain. However, data interchangeability between EHR systems is limited due to the use of different coding standards across systems. Existing methods of…

Computation and Language · Computer Science 2019-05-07 Rohollah Soltani , Alexandre Tomberg

Mutation testing is an effective technique for assessing the effectiveness of test suites by systematically injecting artificial faults into programs. However, existing mutation testing techniques fall short in capturing many types of…

Software Engineering · Computer Science 2026-01-28 Saba Alimadadi , Golnaz Gharachorlu

Cluster analyses of high-dimensional data are often hampered by the presence of large numbers of variables that do not provide relevant information, as well as the perennial issue of choosing an appropriate number of clusters. These…

Computation · Statistics 2024-12-02 Emma Prevot , Rory Toogood , Filippo Pagani , Paul D. W. Kirk

Accessing sensitive patient data for machine learning is challenging due to privacy concerns. Datasets with annotations of personally identifiable information are crucial for developing and testing anonymization systems to enable safe data…

Computation and Language · Computer Science 2026-03-17 Ibrahim Baroud , Christoph Otto , Vera Czehmann , Christine Hovhannisyan , Lisa Raithel , Sebastian Möller , Roland Roller

Vibrational frequency calculations performed under the harmonic approximation are widespread across chemistry. However, it is well-known that the calculated harmonic frequencies tend to systematically overestimate experimental fundamental…

Chemical Physics · Physics 2021-10-27 Juan C. Zapata Trujillo , Laura K. McKemmish

Human cancers present a significant public health challenge and require the discovery of novel drugs through translational research. Transcriptomics profiling data that describes molecular activities in tumors and cancer cell lines are…

High throughput sequencing is a technology that allows for the generation of millions of reads of genomic data regarding a study of interest, and data from high throughput sequencing platforms are usually count compositions. Subsequent…

Quantitative Methods · Quantitative Biology 2017-04-07 Jia R. Wu , Jean M. Macklaim , Briana L. Genge , Gregory B. Gloor

There is a growing need for unbiased clustering methods, ideally automated. We have developed a topology-based analysis tool called Two-Tier Mapper (TTMap) to detect subgroups in global gene expression datasets and identify their…

Genomics · Quantitative Biology 2018-01-08 Rachel Jeitziner , Mathieu Carrière , Jacques Rougemont , Steve Oudot , Kathryn Hess , Cathrin Brisken

Availability of research datasets is keystone for health and life science study reproducibility and scientific progress. Due to the heterogeneity and complexity of these data, a main challenge to be overcome by research data management…

Information Retrieval · Computer Science 2017-09-12 Douglas Teodoro , Luc Mottin , Julien Gobeill , Arnaud Gaudinat , Thérèse Vachon , Patrick Ruch

The typical process for classifying and submitting a newly sequenced virus to the NCBI database involves two steps. First, a BLAST search is performed to determine likely family candidates. That is followed by checking the candidate…

Genomics · Quantitative Biology 2016-03-22 Troy Hernandez , Jie Yang

Automated data visualization plays a crucial role in simplifying data interpretation, enhancing decision-making, and improving efficiency. While large language models (LLMs) have shown promise in generating visualizations from natural…

Computation and Language · Computer Science 2025-07-29 Mizanur Rahman , Md Tahmid Rahman Laskar , Shafiq Joty , Enamul Hoque

Mapper, a topological algorithm, is frequently used as an exploratory tool to build a graphical representation of data. This representation can help to gain a better understanding of the intrinsic shape of high-dimensional genomic data and…

Genomics · Quantitative Biology 2023-07-19 Erik J. Amézquita , Farzana Nasrin , Kathleen M. Storey , Masato Yoshizawa

Benchmark data sets are a cornerstone of machine learning development and applications, ensuring new methods are robust, reliable and competitive. The relative rarity of benchmark sets in computational science, due to the uniqueness of the…

Machine Learning · Computer Science 2025-07-01 Amanda S Barnard