English
Related papers

Related papers: SMAGEXP: a galaxy tool suite for transcriptomics d…

200 papers

Introduction: Feature selection and gene set analysis are of increasing interest in bioinformatics. While these two approaches have been developed for different purposes, we describe how some gene set analysis methods can be used to conduct…

Methodology · Statistics 2015-11-30 Suyan Tian , Chi Wang , Howard H. Chang

Single-cell transcriptomics techniques, such as scRNA-seq, attempt to characterize gene expression profiles in each cell of a heterogeneous sample individually. Due to growing amounts of data generated and the increasing complexity of the…

Genomics · Quantitative Biology 2023-05-02 Laura Puente-Santamaría , Luis del Peso

Accurate sample classification using transcriptomics data is crucial for advancing personalized medicine. Achieving this goal necessitates determining a suitable sample size that ensures adequate statistical power without undue resource…

Methodology · Statistics 2024-09-11 Yunhui Qi , Xinyi Wang , Li-Xuan Qin

This paper investigates methods for improving generative data augmentation for deep learning. Generative data augmentation leverages the synthetic samples produced by generative models as an additional dataset for classification with small…

Machine Learning · Computer Science 2023-10-24 Shin'ya Yamaguchi , Daiki Chijiwa , Sekitoshi Kanai , Atsutoshi Kumagai , Hisashi Kashima

The marine environment is one of the most important sources for microbial biodiversity on the planet. These microbes are drivers for many biogeochemical processes, and their enormous genetic potential is still not fully explored or…

Distributed, Parallel, and Cluster Computing · Computer Science 2016-04-15 Espen Mikal Robertsen , Tim Kahlke , Inge Alexander Raknes , Edvard Pedersen , Erik Kjærner Semb , Martin Ernstsen , Lars Ailo Bongo , Nils Peder Willassen

Synthetic tabular data generation becomes crucial when real data is limited, expensive to collect, or simply cannot be used due to privacy concerns. However, producing good quality synthetic data is challenging. Several probabilistic,…

Machine Learning · Computer Science 2024-06-11 Vikram S Chundawat , Ayush K Tarun , Murari Mandal , Mukund Lahoti , Pratik Narang

Single-cell RNA sequencing (scRNA-seq) reveals cell heterogeneity, with cell clustering playing a key role in identifying cell types and marker genes. Recent advances, especially graph neural networks (GNNs)-based methods, have…

Genomics · Quantitative Biology 2025-10-03 Ping Xu , Zhiyuan Ning , Pengjiang Li , Wenhao Liu , Pengyang Wang , Jiaxu Cui , Yuanchun Zhou , Pengfei Wang

We develop a model-based methodology for integrating gene-set information with an experimentally-derived gene list. The methodology uses a previously reported sampling model, but takes advantage of natural constraints in the…

Methodology · Statistics 2015-06-02 Zhishi Wang , Qiuling He , Bret Larget , Michael A. Newton

In machine learning (ML), ensemble methods such as bagging, boosting, and stacking are widely-established approaches that regularly achieve top-notch predictive performance. Stacking (also called "stacked generalization") is an ensemble…

Machine Learning · Computer Science 2024-04-19 Angelos Chatzimparmpas , Rafael M. Martins , Kostiantyn Kucher , Andreas Kerren

The exponential increase in academic publications has made it increasingly difficult for researchers to remain up to date and systematically synthesize knowledge scattered across vast and fragmented research domains. Literature reviews,…

Digital Libraries · Computer Science 2025-06-12 Kiran Sharmaa , Parul Khurana , Ziya Uddina

This study investigates the automation of meta-analysis in scientific documents using large language models (LLMs). Meta-analysis is a robust statistical method that synthesizes the findings of multiple studies support articles to provide a…

Computation and Language · Computer Science 2024-11-19 Jawad Ibn Ahad , Rafeed Mohammad Sultan , Abraham Kaikobad , Fuad Rahman , Mohammad Ruhul Amin , Nabeel Mohammed , Shafin Rahman

The objective of many high-dimensional microarray and RNA-seq studies is to develop a classifier of cancer patients based on characteristics of their disease. The germinal center B-cell (GCB) classifier study in lymphoma and the National…

Applications · Statistics 2015-09-17 Sandra Safo , Xiao Song , Kevin K. Dobbin

Non-sharable sensitive data collection and analysis in large-scale consortia for genomic research is complicated. Time consuming issues in installing software arise due to different operating systems, software dependencies and running the…

Recent advances in big data and analytics research have provided a wealth of large data sets that are too big to be analyzed in their entirety, due to restrictions on computer memory or storage size. New Bayesian methods have been developed…

Applications · Statistics 2014-09-30 Alexey Miroshnikov , Erin Conlon

Motivation: Networks underlie the generation and interpretation of many biological datasets: gene networks shed light on the regulatory structure of the genome, and cell networks can capture structure of the tumor micro-environment.…

Machine Learning · Statistics 2026-03-18 Bailey Andrew , Erica L. Harris , James A. Poulter , David R. Westhead , Luisa Cutillo

Bayesian inference is on the rise, partly because it allows researchers to quantify parameter uncertainty, evaluate evidence for competing hypotheses, incorporate model ambiguity, and seamlessly update knowledge as information accumulates.…

Methodology · Statistics 2025-09-15 František Bartoš , Eric-Jan Wagenmakers

Empirical claims often rely on one population, design, and analysis. Many-analysts, multiverse, and robustness studies expose how results can vary across plausible analytic choices. Synthesizing these results, however, is nontrivial as all…

Methodology · Statistics 2025-11-24 František Bartoš , Suzanne Hoogeveen , Alexandra Sarafoglou , Samuel Pawel

We introduce a parallel algorithmic architecture for metagenomic sequence assembly, termed MetaPar, which allows for significant reductions in assembly time and consequently enables the processing of large genomic datasets on computers with…

Quantitative Methods · Quantitative Biology 2013-11-18 Minji Kim , Jonathan G. Ligo , Amin Emad , Farzad Farnoud , Olgica Milenkovic , Venugopal V. Veeravalli

The gene set analysis (GSA) is a foundational approach for uncovering the molecular functions associated with a group of genes. Recently, LLM-powered methods have emerged to annotate gene sets with biological functions together with…

Genomics · Quantitative Biology 2025-09-16 Zhizheng Wang , Yifan Yang , Qiao Jin , Zhiyong Lu

The regulAS software package is a bioinformatics tool designed to support computational biology researchers in investigating regulatory mechanisms of splicing alterations through integrative analysis of large-scale RNA-Seq data from cancer…

Genomics · Quantitative Biology 2023-07-26 Sofya Lipnitskaya