English
Related papers

Related papers: Analyzing Chromatin Using Tiled Binned Scatterplot…

200 papers

Missing data is a major challenge in clinical research. In electronic medical records, often a large fraction of the values in laboratory tests and vital signs are missing. The missingness can lead to biased estimates and limit our ability…

Machine Learning · Computer Science 2023-04-18 Omer Noy , Ron Shamir

One of the most important problems in development is how epigenetic domains can be first established, and then maintained, within cells. To address this question, we propose a framework which couples 3D chromatin folding dynamics, to a…

Soft Condensed Matter · Physics 2016-10-25 Davide Michieletto , Enzo Orlandini , Davide Marenduzzo

The discovery of patterns associated with diagnosis, prognosis, and therapy response in digital pathology images often requires intractable labeling of large quantities of histological objects. Here we release an open-source labeling tool,…

In data science, there is a long history of using synthetic data for method development, feature selection and feature engineering. Our current interest in synthetic data comes from recent work in explainability. Today's datasets are…

Machine Learning · Computer Science 2020-07-22 Brian Barr , Ke Xu , Claudio Silva , Enrico Bertini , Robert Reilly , C. Bayan Bruss , Jason D. Wittenbach

Multiple instance learning exhibits a powerful approach for whole slide image-based diagnosis in the absence of pixel- or patch-level annotations. In spite of the huge size of hole slide images, the number of individual slides is often…

Advances in molecular "omics'" technologies have motivated new methodology for the integration of multiple sources of high-content biomedical data. However, most statistical methods for integrating multiple data matrices only consider data…

Machine Learning · Statistics 2020-02-10 Jun Young Park , Eric F. Lock

Finite mixture models are frequently used to uncover latent structures in high-dimensional datasets (e.g.\ identifying clusters of patients in electronic health records). The inference of such structures can be performed in a Bayesian…

Determining the number of factors in high-dimensional factor modeling is essential but challenging, especially when the data are heavy-tailed. In this paper, we introduce a new estimator based on the spectral properties of Spearman sample…

Methodology · Statistics 2024-08-29 Jiaxin Qiu , Zeng Li , Jianfeng Yao

The role of Artificial Intelligence (AI) is growing in every stage of drug development. Nevertheless, a major challenge in drug discovery AI remains: Drug pharmacokinetic (PK) and Drug-Target Interaction (DTI) datasets collected in…

Quantitative Methods · Quantitative Biology 2025-10-27 Bing Hu , Jong-Hoon Park , Helen Chen , Young-Rae Cho , Anita Layton

The proliferation of high-dimensional data from sources such as social media, sensor networks, and online platforms has created new challenges for clustering algorithms. Multi-view clustering, which integrates complementary information from…

Machine Learning · Computer Science 2026-01-23 Chakib Fettal , Lazhar Labiod , Mohamed Nadif

We consider algorithms for going from a "full" matrix to a condensed "band bidiagonal" form using orthogonal transformations. We use the framework of "algorithms by tiles". Within this framework, we study: (i) the tiled bidiagonalization…

Mathematical Software · Computer Science 2016-11-23 Mathieu Faverge , Julien Langou , Yves Robert , Jack Dongarra

In recent years, with the development of microarray technique, discovery of useful knowledge from microarray data has become very important. Biclustering is a very useful data mining technique for discovering genes which have similar…

Computational Engineering, Finance, and Science · Computer Science 2009-09-09 Mohsen lashkargir , S. Amirhassan Monadjemi , Ahmad Baraani Dastjerdi

Data sites selected from modeling high-dimensional problems often appear scattered in non-paternalistic ways. Except for sporadic clustering at some spots, they become relatively far apart as the dimension of the ambient space grows. These…

Numerical Analysis · Mathematics 2021-09-28 Shao-Bo Lin , Xiangyu Chang , Xingping Sun

We present new findings in regard to data analysis in very high dimensional spaces. We use dimensionalities up to around one million. A particular benefit of Correspondence Analysis is its suitability for carrying out an orthonormal…

Machine Learning · Statistics 2015-12-15 Fionn Murtagh

Clinical research often focuses on complex traits in which many variables play a role in mechanisms driving, or curing, diseases. Clinical prediction is hard when data is high-dimensional, but additional information, like domain knowledge…

Methodology · Statistics 2020-05-21 Mirrelijn M. van Nee , Lodewyk F. A. Wessels , Mark A. van de Wiel

Identifying molecular signatures from complex disease patients with underlying symptomatic similarities is a significant challenge in the analysis of high dimensional multi-omics data. Topological data analysis (TDA) provides a way of…

Genomics · Quantitative Biology 2024-04-23 Davide Gurnari , Aldo Guzmán-Sáenz , Filippo Utro , Aritra Bose , Saugata Basu , Laxmi Parida

Point sets in 2D with multiple classes are a common type of data. A canonical visualization design for them are scatterplots, which do not scale to large collections of points. For these larger data sets, binned aggregation (or binning) is…

Human-Computer Interaction · Computer Science 2020-01-16 Florian Heimerl , Chih-Ching Chang , Alper Sarikaya , Michael Gleicher

Ptychography enables coherent diffractive imaging (CDI) of extended samples by raster scanning across the illuminating XUV/X-ray beam thereby generalizing the unique advantages of CDI techniques. Table-top realizations of this method are…

Large-scale biobanks are being collected around the world in efforts to better understand human health and risk factors for disease. They often survey hundreds of thousands of individuals, combining questionnaires with clinical, genetic,…

Quantitative Methods · Quantitative Biology 2019-03-19 Qifan Yang , Gennady V. Roshchupkin , Wiro J. Niessen , Sarah E. Medland , Alyssa H. Zhu , Paul M. Thompson , Neda Jahanshad

The effective visualization of genomic data is crucial for exploring and interpreting complex relationships within and across genes and genomes. Despite advances in developing dedicated bioinformatics software, common visualization tools…

Genomics · Quantitative Biology 2024-11-22 Thomas Hackl , Markus Ankenbrand , Bart van Adrichem , David Wilkins , Kristina Haslinger