English
Related papers

Related papers: Classification of chemical compounds based on the …

200 papers

Predicting monomer reactivity ratios is crucial for controlling monomer sequence distribution in copolymers and their properties. Traditional experimental methods of determining reactivity ratios are time-consuming and resource-intensive,…

Chemical Physics · Physics 2025-12-24 Habibollah Safari , Mona Bavarian

Subgroup identification (clustering) is an important problem in biomedical research. Gene expression profiles are commonly utilized to define subgroups. Longitudinal gene expression profiles might provide additional information on disease…

Methodology · Statistics 2016-09-27 Jiehuan Sun , Jose D. Herazo-Maya , Naftali Kaminski , Hongyu Zhao , Joshua L. Warren

Quality assessments of models in unsupervised learning and clustering verification in particular have been a long-standing problem in the machine learning research. The lack of robust and universally applicable cluster validity scores often…

Machine Learning · Statistics 2018-03-30 Luzie Helfmann , Johannes von Lindheim , Mattes Mollenhauer , Ralf Banisch

Unsupervised Anomaly Detection (UAD) plays a crucial role in identifying abnormal patterns within data without labeled examples, holding significant practical implications across various domains. Although the individual contributions of…

Machine Learning · Computer Science 2024-06-04 Zeyu Fang , Ming Gu , Sheng Zhou , Jiawei Chen , Qiaoyu Tan , Haishuai Wang , Jiajun Bu

Chemical toxicity prediction using machine learning is important in drug development to reduce repeated animal and human testing, thus saving cost and time. It is highly recommended that the predictions of computational toxicology models…

Quantitative Methods · Quantitative Biology 2020-09-28 Kar Wai Lim , Bhanushee Sharma , Payel Das , Vijil Chenthamarakshan , Jonathan S. Dordick

Predicting the effect of mutations in proteins is one of the most critical challenges in protein engineering; by knowing the effect a substitution of one (or several) residues in the protein's sequence has on its overall properties, could…

Computational Engineering, Finance, and Science · Computer Science 2020-10-08 David Medina-Ortiz , Sebastian Contreras , Juan Amado-Hinojosa , Jorge Torres-Almonacid , Juan A. Asenjo , Marcelo Navarrete , Álvaro Olivera-Nappa

Recently, several classifiers that combine primary tumor data, like gene expression data, and secondary data sources, such as protein-protein interaction networks, have been proposed for predicting outcome in breast cancer. In these…

Machine Learning · Computer Science 2013-10-15 C. Staiger , S. Cadot , R. Kooter , M. Dittrich , T. Mueller , G. W. Klau , L. F. A. Wessels

Background: When planning a cluster randomized trial, evaluators often have access to an enumerated cohort representing the target population of clusters. Practicalities of conducting the trial, such as the need to oversample clusters with…

Methodology · Statistics 2024-09-19 Sarah E. Robertson , Jon A. Steingrimsson , Issa J. Dahabreh

Genomic signal processing has been used successfully in bioinformatics to analyze biomolecular sequences and gain varied insights into DNA structure, gene organization, protein binding, sequence evolution, etc. But challenges remain in…

Genomics · Quantitative Biology 2022-11-04 Saish Jaiswal , Shreya Nema , Hema A Murthy , Manikandan Narayanan

The identification of sets of co-regulated genes that share a common function is a key question of modern genomics. Bayesian profile regression is a semi-supervised mixture modelling approach that makes use of a response to guide inference…

We consider how to generate chemical reaction networks (CRNs) from functional specifications. We propose a two-stage approach that combines synthesis by satisfiability modulo theories and Markov chain Monte Carlo based optimisation. First,…

Emerging Technologies · Computer Science 2015-08-19 Neil Dalchau , Niall Murphy , Rasmus Petersen , Boyan Yordanov

This paper explores the use of clustering methods and machine learning algorithms, including Natural Language Processing (NLP), to identify and classify problems identified in credit risk models through textual information contained in…

Machine Learning · Computer Science 2023-06-05 Szymon Lis , Mariusz Kubkowski , Olimpia Borkowska , Dobromił Serwa , Jarosław Kurpanik

We present and review Coupled Two Way Clustering, a method designed to mine gene expression data. The method identifies submatrices of the total expression matrix, whose clustering analysis reveals partitions of samples (and genes) into…

Biological Physics · Physics 2007-05-23 Gad Getz , Hilah Gal , Itai Kela , Eytan Domany , Dan A. Notterman

With the recent growth in data availability and complexity, and the associated outburst of elaborate modelling approaches, model selection tools have become a lifeline, providing objective criteria to deal with this increasingly challenging…

Methodology · Statistics 2020-10-08 Alessandro Casa , Luca Scrucca , Giovanna Menardi

Transcription factors regulate gene expression, but how these proteins recognize and specifically bind to their DNA targets is still debated. Machine learning models are effective means to reveal interaction mechanisms. Here we studied the…

Quantum Physics · Physics 2018-03-02 Richard Y. Li , Rosa Di Felice , Remo Rohs , Daniel A. Lidar

Clustering is a well-known unsupervised machine learning approach capable of automatically grouping discrete sets of instances with similar characteristics. Constrained clustering is a semi-supervised extension to this process that can be…

Machine Learning · Computer Science 2023-03-02 Germán González-Almagro , Daniel Peralta , Eli De Poorter , José-Ramón Cano , Salvador García

In this work we seek clusters of genomic words in human DNA by studying their inter-word lag distributions. Due to the particularly spiked nature of these histograms, a clustering procedure is proposed that first decomposes each…

Applications · Statistics 2021-01-13 Ana Helena Tavares , Jakob Raymaekers , Peter J. Rousseeuw , Paula Brito , Vera Afreixo

In this paper, we address the problem of identifying protein functionality using the information contained in its aminoacid sequence. We propose a method to define sequence similarity relationships that can be used as input for…

Applications · Statistics 2007-11-12 A. G. Flesia , R. Fraiman , F. G. Leonardi

The training of molecular models of quantum mechanical properties based on statistical machine learning requires large datasets which exemplify the map from chemical structure to molecular property. Intelligent a priori selection of…

In this work, we formulated a real-world problem related to sewer pipeline gas detection using the classification-based approaches. The primary goal of this work was to identify the hazardousness of sewer pipeline to offer safe and…

Neural and Evolutionary Computing · Computer Science 2017-07-04 Varun Kumar Ojha , Paramartha Dutta , Atal Chaudhuri