English
Related papers

Related papers: MOD-Finder: Identify multi-omics data sets related…

200 papers

SimOmics is an R package designed to generate realistic, multivariate, and multi-omics synthetic datasets. It is intended for use in benchmarking, method development, and reproducibility in bioinformatics, particularly in the context of…

Genomics · Quantitative Biology 2025-07-15 Kaitao Lai

Motivation: Cancer is heterogeneous, affecting the precise approach to personalized treatment. Accurate subtyping can lead to better survival rates for cancer patients. High-throughput technologies provide multiple omics data for cancer…

Machine Learning · Computer Science 2022-08-01 Hai Yang , Yuhang Sheng , Yi Jiang , Xiaoyang Fang , Dongdong Li , Jing Zhang , Zhe Wang

Multiomics data fusion integrates diverse data modalities, ranging from transcriptomics to proteomics, to gain a comprehensive understanding of biological systems and enhance predictions on outcomes of interest related to disease phenotypes…

Quantitative Methods · Quantitative Biology 2023-08-04 Daisy Yi Ding , Xiaotao Shen , Michael Snyder , Robert Tibshirani

High-throughput technologies to collect field data have made observations possible at scale in several branches of life sciences. The data collected can range from the molecular level (genotypes) to physiological (phenotypic traits) and…

Computational Geometry · Computer Science 2021-07-07 Youjia Zhou , Methun Kamruzzaman , Patrick Schnable , Bala Krishnamoorthy , Ananth Kalyanaraman , Bei Wang

Information extraction from chemistry literature is vital for constructing up-to-date reaction databases for data-driven chemistry. Complete extraction requires combining information across text, tables, and figures, whereas prior work has…

Machine Learning · Computer Science 2024-04-03 Vincent Fan , Yujie Qian , Alex Wang , Amber Wang , Connor W. Coley , Regina Barzilay

Artificial Intelligence predicts drug properties by encoding drug molecules, aiding in the rapid screening of candidates. Different molecular representations, such as SMILES and molecule graphs, contain complementary information for…

Machine Learning · Computer Science 2024-06-27 Muzhen Cai , Sendong Zhao , Haochun Wang , Yanrui Du , Zewen Qiang , Bing Qin , Ting Liu

Many-Objective Feature Selection (MOFS) approaches use four or more objectives to determine the relevance of a subset of features in a supervised learning task. As a consequence, MOFS typically returns a large set of non-dominated…

Machine Learning · Computer Science 2023-12-01 Uchechukwu F. Njoku , Alberto Abelló , Besim Bilalli , Gianluca Bontempi

This paper introduces M$^{3}$-20M, a large-scale Multi-Modal Molecule dataset that contains over 20 million molecules, with the data mainly being integrated from existing databases and partially generated by large language models. Designed…

Quantitative Methods · Quantitative Biology 2025-03-18 Siyuan Guo , Lexuan Wang , Chang Jin , Jinxian Wang , Han Peng , Huayang Shi , Wengen Li , Jihong Guan , Shuigeng Zhou

Understanding the interaction between different drugs (drug-drug interaction or DDI) is critical for ensuring patient safety and optimizing therapeutic outcomes. Existing DDI datasets primarily focus on textual information, overlooking…

Machine Learning · Computer Science 2025-06-03 Tung-Lam Ngo , Ba-Hoang Tran , Duy-Cat Can , Trung-Hieu Do , Oliver Y. Chén , Hoang-Quynh Le

Modeling with multi-omics data presents multiple challenges such as the high-dimensionality of the problem ($p \gg n$), the presence of interactions between features, and the need for integration between multiple data sources. We establish…

Methodology · Statistics 2024-09-17 Matteo D'Alessandro , Theophilus Quachie Asenso , Manuela Zucknick

Existing multi-modal learning methods on fundus and OCT images mostly require both modalities to be available and strictly paired for training and testing, which appears less practical in clinical scenarios. To expand the scope of clinical…

Computer Vision and Pattern Recognition · Computer Science 2025-04-08 Lehan Wang , Chongchong Qi , Chubin Ou , Lin An , Mei Jin , Xiangbin Kong , Xiaomeng Li

The automated extraction of chemical structures and their corresponding bioactivity data is essential for accelerating drug discovery and enabling data-driven research. Current optical chemical structure recognition tools lack the…

Quantitative Methods · Quantitative Biology 2026-03-04 Zhe Wang , Fangtian Fu , Wei Zhang , Lige Yan , Nan Li , Wenxia Deng , Yan Meng , Jianping Wu , Hui Wu , Wenting Wu , Gang Xu , Xiang Li , Si Chen

Muilti-modality data are ubiquitous in biology, especially that we have entered the multi-omics era, when we can measure the same biological object (cell) from different aspects (omics) to provide a more comprehensive insight into the…

Genomics · Quantitative Biology 2021-12-15 Xuesong Wang , Zhihang Hu , Tingyang Yu , Ruijie Wang , Yumeng Wei , Juan Shu , Jianzhu Ma , Yu Li

The extraction of molecular structures and reaction data from scientific documents is challenging due to their varied, unstructured chemical formats and complex document layouts. To address this, we introduce MolMole, a vision-based deep…

Artificial intelligence (AI) has demonstrated significant promise in advancing organic chemistry research; however, its effectiveness depends on the availability of high-quality chemical reaction data. Currently, most published chemical…

Computer Vision and Pattern Recognition · Computer Science 2025-03-12 Yufan Chen , Ching Ting Leung , Jianwei Sun , Yong Huang , Linyan Li , Hao Chen , Hanyu Gao

Researchers and scientists increasingly rely on specialized information retrieval (IR) or recommendation systems (RS) to support them in their daily research tasks. Paper recommender systems are one such tool scientists use to stay on top…

Information Retrieval · Computer Science 2022-05-12 Corinna Breitinger , Kay Herklotz , Tim Flegelskamp , Norman Meuschke

The concept of personalised medicine in cancer therapy is becoming increasingly important. There already exist drugs administered specifically for patients with tumours presenting well-defined mutations. However, the field is still in its…

Biomolecules · Quantitative Biology 2024-08-26 Abbi Abdel-Rehim , Oghenejokpeme Orhobor , Gareth Griffiths , Larisa Soldatova , Ross D. King

The increasing quantity of multi-omics data, such as methylomic and transcriptomic profiles, collected on the same specimen, or even on the same cell, provide a unique opportunity to explore the complex interactions that define cell…

Bump-hunting or mode identification is a fundamental problem that arises in almost every scientific field of data-driven discovery. Surprisingly, very few data modeling tools are available for automatic (not requiring manual case-by-base…

Methodology · Statistics 2016-11-10 Subhadeep Mukhopadhyay

To fully expedite AI-powered chemical research, high-quality chemical databases are the foundation. Automatic extraction of chemical information from the literature is essential for constructing reaction databases, but it is currently…

Artificial Intelligence · Computer Science 2026-03-09 Yufan Chen , Ching Ting Leung , Bowen Yu , Jianwei Sun , Yong Huang , Linyan Li , Hao Chen , Hanyu Gao