中文
相关论文

相关论文: BioUNER: A Benchmark Dataset for Clinical Urdu Nam…

200 篇论文

OCR algorithms have received a significant improvement in performance recently, mainly due to the increase in the capabilities of artificial intelligence algorithms. However, this advancement is not evenly distributed over all languages.…

计算机视觉与模式识别 · 计算机科学 2020-05-15 Atique Ur Rehman , Sibt Ul Hussain

Nowadays, many Natural Language Processing (NLP) tasks see the demand for incorporating knowledge external to the local information to further improve the performance. However, there is little related work on Named Entity Recognition (NER),…

计算与语言 · 计算机科学 2023-03-07 Chiao-Wei Hsu , Keh-Yih Su

A crucial step within secondary analysis of electronic health records (EHRs) is to identify the patient cohort under investigation. While EHRs contain medical billing codes that aim to represent the conditions and treatments patients may…

Clinical data in hospitals are increasingly accessible for research through clinical data warehouses. However these documents are unstructured and it is therefore necessary to extract information from medical reports to conduct clinical…

计算与语言 · 计算机科学 2024-04-04 Rian Touchent , Laurent Romary , Eric de la Clergerie

We introduce KyrgyzNER, the first manually annotated named entity recognition dataset for the Kyrgyz language. Comprising 1,499 news articles from the 24.KG news portal, the dataset contains 10,900 sentences and 39,075 entity mentions…

计算与语言 · 计算机科学 2025-09-24 Timur Turatali , Anton Alekseev , Gulira Jumalieva , Gulnara Kabaeva , Sergey Nikolenko

This paper introduces UQA, a novel dataset for question answering and text comprehension in Urdu, a low-resource language with over 70 million native speakers. UQA is generated by translating the Stanford Question Answering Dataset…

计算与语言 · 计算机科学 2024-07-24 Samee Arif , Sualeha Farid , Awais Athar , Agha Ali Raza

Named Entity Recognition (NER) plays a pivotal role in medical Natural Language Processing (NLP). Yet, there has not been an open-source medical NER dataset specifically for the Korean language. To address this, we utilized ChatGPT to…

计算与语言 · 计算机科学 2024-03-26 Sungjoo Byun , Jiseung Hong , Sumin Park , Dongjun Jang , Jean Seo , Minseok Kim , Chaeyoung Oh , Hyopil Shin

This paper conducts a comprehensive investigation into applying large language models, particularly on BioBERT, in healthcare. It begins with thoroughly examining previous natural language processing (NLP) approaches in healthcare, shedding…

人工智能 · 计算机科学 2023-10-13 Shyni Sharaf , V. S. Anoop

Electronic health records (EHR) contain large volumes of unstructured text, requiring the application of Information Extraction (IE) technologies to enable clinical analysis. We present the open-source Medical Concept Annotation Toolkit…

Large Language Models (LLMs) have recently gained attention in the life sciences due to their capacity to model, extract, and apply complex biological information. Beyond their classical use as chatbots, these systems are increasingly used…

计算与语言 · 计算机科学 2025-07-03 Baqer M. Merzah , Tania Taami , Salman Asoudeh , Saeed Mirzaee , Amir reza Hossein pour , Amir Ali Bengari

Electronic health records (EHRs) form an invaluable resource for training clinical decision support systems. To leverage the potential of such systems in high-risk applications, we need large, structured tabular datasets on which we can…

人工智能 · 计算机科学 2025-11-24 Paloma Rabaey , Adrick Tench , Stefan Heytens , Thomas Demeester

As more and more Arabic texts emerged on the Internet, extracting important information from these Arabic texts is especially useful. As a fundamental technology, Named entity recognition (NER) serves as the core component in information…

计算与语言 · 计算机科学 2023-08-09 Xiaoye Qu , Yingjie Gu , Qingrong Xia , Zechang Li , Zhefeng Wang , Baoxing Huai

In this paper, we introduce the NER dataset from CLUE organization (CLUENER2020), a well-defined fine-grained dataset for named entity recognition in Chinese. CLUENER2020 contains 10 categories. Apart from common labels like person,…

计算与语言 · 计算机科学 2020-01-22 Liang Xu , Yu tong , Qianqian Dong , Yixuan Liao , Cong Yu , Yin Tian , Weitang Liu , Lu Li , Caiquan Liu , Xuanwei Zhang

As a major social media platform, Twitter publishes a large number of user-generated text (tweets) on a daily basis. Mining such data can be used to address important social, public health, and emergency management issues that are…

计算与语言 · 计算机科学 2021-12-07 Qing Han , Shubo Tian , Jinfeng Zhang

Disease name recognition and normalization, which is generally called biomedical entity linking, is a fundamental process in biomedical text mining. Recently, neural joint learning of both tasks has been proposed to utilize the mutual…

计算与语言 · 计算机科学 2021-04-22 Shogo Ujiie , Hayate Iso , Shuntaro Yada , Shoko Wakamiya , Eiji Aramaki

Deep learning models (aka Deep Neural Networks) have revolutionized many fields including computer vision, natural language processing, speech recognition, and is being increasingly used in clinical healthcare applications. However, few…

机器学习 · 计算机科学 2017-10-25 Sanjay Purushotham , Chuizheng Meng , Zhengping Che , Yan Liu

In this paper, we propose DiffusionNER, which formulates the named entity recognition task as a boundary-denoising diffusion process and thus generates named entities from noisy spans. During training, DiffusionNER gradually adds noises to…

计算与语言 · 计算机科学 2023-05-23 Yongliang Shen , Kaitao Song , Xu Tan , Dongsheng Li , Weiming Lu , Yueting Zhuang

With the ever-growing popularity of the field of NLP, the demand for datasets in low resourced-languages follows suit. Following a previously established framework, in this paper, we present the UNER dataset, a multilingual and hierarchical…

计算与语言 · 计算机科学 2022-12-16 Diego Alves , Gaurish Thakkar , Gabriel Amaral , Tin Kuculo , Marko Tadić

Small and imbalanced datasets commonly seen in healthcare represent a challenge when training classifiers based on deep learning models. So motivated, we propose a novel framework based on BioBERT (Bidirectional Encoder Representations from…

计算与语言 · 计算机科学 2020-06-23 Shijing Si , Rui Wang , Jedrek Wosik , Hao Zhang , David Dov , Guoyin Wang , Ricardo Henao , Lawrence Carin

Automatically extracting organization names from the affiliation sentences of articles related to biomedicine is of great interest to the pharmaceutical marketing industry, health care funding agencies and public health officials. It will…

数字图书馆 · 计算机科学 2010-05-17 Siddhartha Jonnalagadda , Philip Topham , Graciela Gonzalez
‹ 上一页 1 8 9 10 下一页 ›