English
Related papers

Related papers: A large dataset curation and benchmark for drug ta…

200 papers

Sharing protected health information (PHI) is critical for furthering biomedical research. Before data can be distributed, practitioners often perform deidentification to remove any PHI contained in the text. Contemporary deidentification…

Computation and Language · Computer Science 2024-10-23 John X. Morris , Thomas R. Campion , Sri Laasya Nutheti , Yifan Peng , Akhil Raj , Ramin Zabih , Curtis L. Cole

With the rapid development of high-throughput technologies, parallel acquisition of large-scale drug-informatics data provides huge opportunities to improve pharmaceutical research and development. One significant application is the purpose…

Machine Learning · Computer Science 2018-10-03 Lingwei Xie , Song He , Shu Yang , Boyuan Feng , Kun Wan , Zhongnan Zhang , Xiaochen Bo , Yufei Ding

Modern self-driving autonomy systems heavily rely on deep learning. As a consequence, their performance is influenced significantly by the quality and richness of the training data. Data collecting platforms can generate many hours of raw…

Machine Learning · Computer Science 2021-01-19 Abbas Sadat , Sean Segal , Sergio Casas , James Tu , Bin Yang , Raquel Urtasun , Ersin Yumer

We explore the hyperparameters and introduce a methodological framework to convert disease patterns from time series data of blood test results into correlation graphs for causal hypothesis exploration. The networks represent hypotheses…

Other Quantitative Biology · Quantitative Biology 2025-12-29 David Patrick Duys Montealegre , Alexander Fulton , Mahta Haghighat Ghahfarokhi , Abicumaran Uthamacumaran , Hector Zenil

Data curation is a field with origins in librarianship and archives, whose scholarship and thinking on data issues go back centuries, if not millennia. The field of machine learning is increasingly observing the importance of data curation…

Computers and Society · Computer Science 2025-01-06 Eshta Bhardwaj , Harshit Gujral , Siyi Wu , Ciara Zogheib , Tegan Maharaj , Christoph Becker

This study investigates the identification power gained by combining experimental data, in which treatment is randomized, with observational data, in which treatment is self-selected, for distributional treatment effect (DTE) parameters.…

Econometrics · Economics 2026-04-24 Shosei Sakaguchi

New technologies and equipment allow for mass treatment of samples and research teams share acquired data on an always larger scale. In this context scientists are facing a major data exploitation problem. More precisely, using these data…

Quantitative Methods · Quantitative Biology 2009-07-02 Julie Bourbeillon , Catherine Garbay , Françoise Giroud

The discovery of clinical biomarkers requires large patient cohorts and is aided by a pooled data approach across institutions. In many countries, data protection constraints, especially in the clinical environment, forbid the exchange of…

Machine Learning · Statistics 2023-10-03 Daniela Zöller , Harald Binder

Motivation: Predicting the drug-target interaction is crucial for drug discovery as well as drug repurposing. Machine learning is commonly used in drug-target affinity (DTA) problem. However, machine learning model faces the cold-start…

Biomolecules · Quantitative Biology 2022-02-03 Tri Minh Nguyen , Thin Nguyen , Truyen Tran

Clinical notes contain an abundance of important but not-readily accessible information about patients. Systems to automatically extract this information rely on large amounts of training data for which their exists limited resources to…

Computation and Language · Computer Science 2020-04-23 Andriy Mulyar , Bridget T. McInnes

Estimating heterogeneous treatment effects is central to data-driven decision-making, yet industrial applications often face a fundamental tension between limited randomized controlled trial (RCT) budgets and abundant but biased…

Recommendation systems must continuously adapt to evolving user behavior, yet the volume of data generated in large-scale streaming environments makes frequent full retraining impractical. This work investigates how targeted data selection…

Bioinformatics research is characterized by voluminous and incremental datasets and complex data analytics methods. The machine learning methods used in bioinformatics are iterative and parallel. These methods can be scaled to handle big…

Computational Engineering, Finance, and Science · Computer Science 2015-06-17 Hirak Kashyap , Hasin Afzal Ahmed , Nazrul Hoque , Swarup Roy , Dhruba Kumar Bhattacharyya

Several techniques have been proposed to address the problem of recognizing activities of daily living from signals. Deep learning techniques applied to inertial signals have proven to be effective, achieving significant classification…

Signal Processing · Electrical Eng. & Systems 2022-01-21 Hamza Amrani , Daniela Micucci , Marco Mobilio , Paolo Napoletano

Protein-protein interactions (PPIs) are critical to normal cellular function and are related to many disease pathways. However, only 4% of PPIs are annotated with PTMs in biological knowledge databases such as IntAct, mainly performed…

Machine Learning · Computer Science 2022-01-10 Aparna Elangovan , Yuan Li , Douglas E. V. Pires , Melissa J. Davis , Karin Verspoor

Computational models that accurately predict the binding affinity of an input protein-chemical pair can accelerate drug discovery studies. These models are trained on available protein-chemical interaction datasets, which may contain…

Quantitative Methods · Quantitative Biology 2023-01-10 Rıza Özçelik , Alperen Bağ , Berk Atıl , Melih Barsbey , Arzucan Özgür , Elif Özkırımlı

Much recent research aims to identify evidence for Drug-Drug Interactions (DDI) and Adverse Drug reactions (ADR) from the biomedical scientific literature. In addition to this "Bibliome", the universe of social media provides a very…

Social and Information Networks · Computer Science 2016-01-15 Rion Brattig Correia , Lang Li , Luis M. Rocha

Diabetes is a chronic disease with a significant global health burden, requiring multi-stakeholder collaboration for optimal management. Large language models (LLMs) have shown promise in various healthcare scenarios, but their…

Computation and Language · Computer Science 2025-03-14 Lai Wei , Zhen Ying , Muyang He , Yutong Chen , Qian Yang , Yanzhe Hong , Jiaping Lu , Kaipeng Zheng , Shaoting Zhang , Xiaoying Li , Weiran Huang , Ying Chen

The task of drug-target interaction prediction holds significant importance in pharmacology and therapeutic drug design. In this paper, we present FRnet-DTI, an auto encoder and a convolutional classifier for feature manipulation and drug…

Machine Learning · Computer Science 2020-05-26 Farshid Rayhan , Sajid Ahmed , Zaynab Mousavian , Dewan Md Farid , Swakkhar Shatabda

The quality of datasets plays an increasingly crucial role in the research and development of modern artificial intelligence (AI). Despite the proliferation of open dataset platforms nowadays, data quality issues, such as incomplete…

Artificial Intelligence · Computer Science 2025-05-28 Benhao Huang , Yingzhuo Yu , Jin Huang , Xingjian Zhang , Jiaqi Ma