English
Related papers

Related papers: Characterizing Diseases from Unstructured Text: A …

200 papers

``Classical'' word embeddings, such as Word2Vec, have been shown to capture the semantics of words based on their distributional properties. However, their ability to represent the different meanings that a word may have is limited. Such…

Computation and Language · Computer Science 2020-04-20 Lea Dieudonat , Kelvin Han , Phyllicia Leavitt , Esteban Marquer

Early detection of preventable diseases is important for better disease management, improved inter-ventions, and more efficient health-care resource allocation. Various machine learning approacheshave been developed to utilize information…

Machine Learning · Computer Science 2018-08-16 Jingshu Liu , Zachariah Zhang , Narges Razavian

Distributional semantics models are known to struggle with small data. It is generally accepted that in order to learn 'a good vector' for a word, a model must have sufficient examples of its usage. This contradicts the fact that humans can…

Computation and Language · Computer Science 2017-07-21 Aurelie Herbelot , Marco Baroni

Diagnostic Captioning (DC) automatically generates a diagnostic text from one or more medical images (e.g., X-rays, MRIs) of a patient. Treated as a draft, the generated text may assist clinicians, by providing an initial estimation of the…

Artificial Intelligence · Computer Science 2024-06-21 Panagiotis Kaliosis , John Pavlopoulos , Foivos Charalampakos , Georgios Moschovis , Ion Androutsopoulos

In the context of personalized medicine, text mining methods pose an interesting option for identifying disease-gene associations, as they can be used to generate novel links between diseases and genes which may complement knowledge from…

Computation and Language · Computer Science 2017-09-28 Hendrik ter Horst , Matthias Hartung , Roman Klinger , Matthias Zwick , Philipp Cimiano

International Classification of Diseases (ICD) are the de facto codes used globally for clinical coding. These codes enable healthcare providers to claim reimbursement and facilitate efficient storage and retrieval of diagnostic…

Computation and Language · Computer Science 2022-02-22 Pavithra Rajendran , Alexandros Zenonos , Josh Spear , Rebecca Pope

In recent years social and news media have increasingly been used to explain patterns in disease activity and progression. Social media data, principally from the Twitter network, has been shown to correlate well with official disease case…

Social and Information Networks · Computer Science 2015-04-17 Donal Simmie , Nicholas Thapen , Chris Hankin

Despite diverse efforts to mine various modalities of medical data, the conversations between physicians and patients at the time of care remain an untapped source of insights. In this paper, we leverage this data to extract structured…

Machine Learning · Computer Science 2020-07-15 Kundan Krishna , Amy Pavel , Benjamin Schloss , Jeffrey P. Bigham , Zachary C. Lipton

We explore the potential of a popular distributional semantics vector space model, word2vec, for capturing meaningful relationships in ecological (complex polyphonic) music. More precisely, the skip-gram version of word2vec is used to model…

Sound · Computer Science 2018-12-03 Ching-Hua Chuan , Kat Agres , Dorien Herremans

Biomedical association studies are increasingly done using clinical concepts, and in particular diagnostic codes from clinical data repositories as phenotypes. Clinical concepts can be represented in a meaningful, vector space using word…

Quantitative Methods · Quantitative Biology 2018-11-06 Brett K. Beaulieu-Jones , Isaac S. Kohane , Andrew L. Beam

Clinical trials are central to medical progress because they help improve understanding of human health and the healthcare system. They play a key role in discovering new ways to detect, prevent, or treat diseases, and it is essential that…

Computation and Language · Computer Science 2025-10-16 Surya Tejaswi Yerramsetty , Almas Fathimah

Electronic Healthcare records contain large volumes of unstructured data in different forms. Free text constitutes a large portion of such data, yet this source of richly detailed information often remains under-used in practice because of…

Computation and Language · Computer Science 2019-10-17 M. Tarik Altuncu , Erik Mayer , Sophia N. Yaliraki , Mauricio Barahona

Motivation: Ontologies are widely used in biology for data annotation, integration, and analysis. In addition to formally structured axioms, ontologies contain meta-data in the form of annotation axioms which provide valuable pieces of…

Computation and Language · Computer Science 2018-05-01 Fatima Zohra Smaili , Xin Gao , Robert Hoehndorf

Causal inference, a critical tool for informing business decisions, traditionally relies heavily on structured data. However, in many real-world scenarios, such data can be incomplete or unavailable. This paper presents a framework that…

Machine Learning · Computer Science 2026-02-17 Boning Zhou , Ziyu Wang , Han Hong , Haoqi Hu

Key features of mental illnesses are reflected in speech. Our research focuses on designing a multimodal deep learning structure that automatically extracts salient features from recorded speech samples for predicting various mental…

Machine Learning · Computer Science 2020-04-15 Habibeh Naderi , Behrouz Haji Soleimani , Stan Matwin

Objective: We aim to learn potential novel cures for diseases from unstructured text sources. More specifically, we seek to extract drug-disease pairs of potential cures to diseases by a simple reasoning over the structure of spoken text.…

Information Retrieval · Computer Science 2020-11-17 Rahul Yedida , Saad Mohammad Abrar , Cleber Melo-Filho , Eugene Muratov , Rada Chirkova , Alexander Tropsha

Effective communication between healthcare providers and patients is crucial to providing high-quality patient care. In this work, we investigate how Doctor-written and AI-generated texts in healthcare consultations can be classified using…

Computation and Language · Computer Science 2024-02-08 Olumide Ebenezer Ojo , Olaronke Oluwayemisi Adebanji , Alexander Gelbukh , Hiram Calvo , Anna Feldman

Word feature vectors have been proven to improve many NLP tasks. With recent advances in unsupervised learning of these feature vectors, it became possible to train it with much more data, which also resulted in better quality of learned…

Computation and Language · Computer Science 2022-11-29 Marius Sajgalik , Michal Barla , Maria Bielikova

A key component of deep learning (DL) for natural language processing (NLP) is word embeddings. Word embeddings that effectively capture the meaning and context of the word that they represent can significantly improve the performance of…

With a simple architecture and the ability to learn meaningful word embeddings efficiently from texts containing billions of words, word2vec remains one of the most popular neural language models used today. However, as only a single…

Machine Learning · Statistics 2017-06-09 Franziska Horn