English
Related papers

Related papers: Finding the centre: corrections for asymmetry in h…

200 papers

Cancer is a complex disease driven by genomic alterations, and tumor sequencing is becoming a mainstay of clinical care for cancer patients. The emergence of multi-institution sequencing data presents a powerful resource for learning…

Genomics · Quantitative Biology 2024-10-31 Yuan Chen , Ronglai Shen , Xiwen Feng , Katherine Panageas

High-throughput data analyses are becoming common in biology, communications, economics and sociology. The vast amounts of data are usually represented in the form of matrices and can be considered as knowledge networks. Spectra-based…

Quantitative Methods · Quantitative Biology 2010-01-06 Viet-Anh Nguyen , Zdena Koukolikova-Nicola , Franco Bagnoli , Pietro Lio

Very long and noisy sequence data arise from biological sciences to social science including high throughput data in genomics and stock prices in econometrics. Often such data are collected in order to identify and understand shifts in…

Methodology · Statistics 2016-07-15 Yue S. Niu , Ning Hao , Heping Zhang

Gene set analysis, a popular approach for analyzing high-throughput gene expression data, aims to identify sets of related genes that show significantly enriched or depleted expression patterns between different conditions. In the last…

RNA-Seq technology allows for studying the transcriptional state of the cell at an unprecedented level of detail. Beyond quantification of whole-gene expression, it is now possible to disentangle the abundance of individual alternatively…

Genomics · Quantitative Biology 2012-10-11 Barbara Rakitsch , Christoph Lippert , Hande Topa , Karsten Borgwardt , Antti Honkela , Oliver Stegle

Recent developments in extracting and processing biological and clinical data are allowing quantitative approaches to studying living systems. High-throughput sequencing, expression profiles, proteomics, and electronic health records are…

Quantitative Methods · Quantitative Biology 2010-10-22 Vladimir Trifonov , Laura Pasqualucci , Riccardo Dalla-Favera , Raul Rabadan

Just as in eukaryotes, high-throughput chromosome conformation capture (Hi-C) data have revealed nested organizations of bacterial chromosomes into overlapping interaction domains. In this chapter, we present a multiscale analysis framework…

Genomics · Quantitative Biology 2020-10-06 Nelle Varoquaux , Virginia S. Lioy , Frédéric Boccard , Ivan Junier

Integrating heterogeneous datasets across different measurement platforms is a fundamental challenge in many scientific applications. A common example arises in deconvolution problems, such as cell type deconvolution, where one aims to…

Methodology · Statistics 2025-09-30 Dongyue Xie , Lin Gui , Jingshu Wang

Although there is no shortage of clustering algorithms proposed in the literature, the question of the most relevant strategy for clustering compositional data (i.e., data made up of profiles, whose rows belong to the simplex) remains…

Statistics Theory · Mathematics 2018-05-17 Antoine Godichon-Baggioni , Cathy Maugis-Rabusseau , Andrea Rau

Research increasingly relies on computational methods to analyze experimental data and predict molecular properties. Current approaches often require researchers to use a variety of tools for statistical analysis and machine learning,…

Quantitative Methods · Quantitative Biology 2025-12-01 Luke Rimmo Lego , Samantha Gauthier , Denver Jn. Baptiste

Although bulk transcriptomic analyses have significantly contributed to an enhanced comprehension of multifaceted diseases, their exploration capacity is impeded by the heterogeneous compositions of biological samples. Indeed, by averaging…

Quantitative Methods · Quantitative Biology 2023-10-24 Bastien Chassagnol , Grégory Nuel , Etienne Becht

An applied problem facing all areas of data science is harmonizing data sources. Joining data from multiple origins with unmapped and only partially overlapping features is a prerequisite to developing and testing robust, generalizable…

A vast array of transformative technologies developed over the past decade has enabled measurement and perturbation at ever increasing scale, yet our understanding of many systems remains limited by experimental capacity. Overcoming this…

Quantitative Methods · Quantitative Biology 2020-12-25 Brian Cleary , Aviv Regev

Datasets are rarely a realistic approximation of the target population. Say, prevalence is misrepresented, image quality is above clinical standards, etc. This mismatch is known as sampling bias. Sampling biases are a major hindrance for…

Technological advances in the past decade, hardware and software alike, have made access to high-performance computing (HPC) easier than ever. We review these advances from a statistical computing perspective. Cloud computing makes access…

Computation · Statistics 2021-07-19 Seyoon Ko , Hua Zhou , Jin J. Zhou , Joong-Ho Won

High throughput screening of compounds (chemicals) is an essential part of drug discovery [7], involving thousands to millions of compounds, with the purpose of identifying candidate hits. Most statistical tools, including the industry…

Machine Learning · Statistics 2017-09-29 Ivo D. Shterev , David B. Dunson , Cliburn Chan , Gregory D. Sempowski

A simple and fast analysis method to sort large data sets into groups with shared distinguishing characteristics is described, and applied to single molecular break junction conductance versus electrode displacement data. The method, based…

Mesoscale and Nanoscale Physics · Physics 2018-01-10 J. M. Hamill , X. T. Zhao , G. Mészáros , M. R. Bryce , M. Arenz

High-dimensional data acquired from biological experiments such as next generation sequencing are subject to a number of confounding effects. These effects include both technical effects, such as variation across batches from instrument…

Machine Learning · Computer Science 2018-12-11 Kabir Manghnani , Adam Drake , Nathan Wan , Imran Haque

Many modern biological assays, including RNA sequencing, yield integer-valued counts that reflect the number of molecules detected. These measurements are often not at the desired resolution: while the unit of interest is typically a single…

Machine Learning · Computer Science 2026-03-06 Nic Fishman , Gokul Gowri , Tanush Kumar , Jiaqi Lu , Valentin de Bortoli , Jonathan S. Gootenberg , Omar Abudayyeh

High-dimensional data of discrete and skewed nature is commonly encountered in high-throughput sequencing studies. Analyzing the network itself or the interplay between genes in this type of data continues to present many challenges. As…

Methodology · Statistics 2017-12-01 Anjali Silva , Steven J. Rothstein , Paul D. McNicholas , Sanjeena Subedi