中文
相关论文

相关论文: A Survey on Awesome Korean NLP Datasets

200 篇论文

Natural Language Processing (NLP) is today a very active field of research and innovation. Many applications need however big sets of data for supervised learning, suitably labelled for the training purpose. This includes applications for…

计算与语言 · 计算机科学 2021-02-23 ElMehdi Boujou , Hamza Chataoui , Abdellah El Mekki , Saad Benjelloun , Ikram Chairi , Ismail Berrada

As LLMs are increasingly deployed in real-world interactions, their social reasoning in interpersonal communication becomes critical. To explore their capabilities, we introduce SCRIPTS, a 1.1k-dialogue dataset in English and Korean,…

计算与语言 · 计算机科学 2026-04-21 Eunsu Kim , Junyeong Park , Juhyun Oh , Kiwoong Park , Seyoung Song , A. Seza Doğruöz , Alice Oh , Najoung Kim

Structured information extraction from scientific literature is crucial for capturing core concepts and emerging trends in specialized fields. While existing datasets aid model development, most focus on specific publication sections due to…

计算与语言 · 计算机科学 2026-04-06 Decheng Duan , Yingyi Zhang , Jitong Peng , Chengzhi Zhang

The necessity of language-specific tokenizers intuitively appears crucial for effective natural language processing, yet empirical analyses on their significance and underlying reasons are lacking. This study explores how language-specific…

计算与语言 · 计算机科学 2025-02-24 Jean Seo , Jaeyoon Kim , SungJoo Byun , Hyopil Shin

Developing documentation guidelines and easy-to-use templates for datasets and models is a challenging task, especially given the variety of backgrounds, skills, and incentives of the people involved in the building of natural language…

The field of Natural Language Processing (NLP) has seen significant advancements with the development of Large Language Models (LLMs). However, much of this research remains focused on English, often overlooking low-resource languages like…

计算与语言 · 计算机科学 2024-08-22 Anh-Dung Vo , Minseong Jung , Wonbeen Lee , Daewoo Choi

The development of Large Language Models (LLMs) requires robust benchmarks that encompass not only academic domains but also industrial fields to effectively evaluate their applicability in real-world scenarios. In this paper, we introduce…

计算与语言 · 计算机科学 2025-07-21 Seokhee Hong , Sunkyoung Kim , Guijin Son , Soyeon Kim , Yeonjung Hong , Jinsik Lee

Dockerfiles are one of the most prevalent kinds of DevOps artifacts used in industry. Despite their prevalence, there is a lack of sophisticated semantics-aware static analysis of Dockerfiles. In this paper, we introduce a dataset of…

软件工程 · 计算机科学 2020-03-31 Jordan Henkel , Christian Bird , Shuvendu K. Lahiri , Thomas Reps

We introduce KoBALT (Korean Benchmark for Advanced Linguistic Tasks), a comprehensive linguistically-motivated benchmark comprising 700 multiple-choice questions spanning 24 phenomena across five linguistic domains: syntax, semantics,…

Despite the recent popularity of knowledge graph (KG) related tasks and benchmarks such as KG embeddings, link prediction, entity alignment and evaluation of the reasoning abilities of pretrained language models as KGs, the structure and…

机器学习 · 计算机科学 2023-11-14 Nedelina Teneva , Estevam Hruschka

Peer review constitutes a core component of scholarly publishing; yet it demands substantial expertise and training, and is susceptible to errors and biases. Various applications of NLP for peer reviewing assistance aim to support reviewers…

计算与语言 · 计算机科学 2023-05-22 Nils Dycke , Ilia Kuznetsov , Iryna Gurevych

We introduce \texttt{N-LTP}, an open-source neural language technology platform supporting six fundamental Chinese NLP tasks: {lexical analysis} (Chinese word segmentation, part-of-speech tagging, and named entity recognition), {syntactic…

计算与语言 · 计算机科学 2021-09-24 Wanxiang Che , Yunlong Feng , Libo Qin , Ting Liu

With the ever-growing amounts of textual data from a large variety of languages, domains, and genres, it has become standard to evaluate NLP algorithms on multiple datasets in order to ensure consistent performance across heterogeneous…

计算与语言 · 计算机科学 2017-09-28 Rotem Dror , Gili Baumer , Marina Bogomolov , Roi Reichart

Named Entity Recognition (NER) plays a pivotal role in medical Natural Language Processing (NLP). Yet, there has not been an open-source medical NER dataset specifically for the Korean language. To address this, we utilized ChatGPT to…

计算与语言 · 计算机科学 2024-03-26 Sungjoo Byun , Jiseung Hong , Sumin Park , Dongjun Jang , Jean Seo , Minseok Kim , Chaeyoung Oh , Hyopil Shin

In pace with developments in the research field of artificial intelligence, knowledge graphs (KGs) have attracted a surge of interest from both academia and industry. As a representation of semantic relations between entities, KGs have…

计算与语言 · 计算机科学 2022-10-04 Phillip Schneider , Tim Schopf , Juraj Vladika , Mikhail Galkin , Elena Simperl , Florian Matthes

Language models have been foundations in various scenarios of NLP applications, but it has not been well applied in language variety studies, even for the most popular language like English. This paper represents one of the few initial…

计算与语言 · 计算机科学 2023-10-10 Yang Liu , Melissa Xiaohui Qin , Long Wang , Chao Huang

The reliance on translated or adapted datasets from English or multilingual resources introduces challenges regarding linguistic and cultural suitability. This study addresses the need for robust and culturally appropriate benchmarks by…

Research in natural language processing (NLP) for Computational Social Science (CSS) heavily relies on data from social media platforms. This data plays a crucial role in the development of models for analysing socio-linguistic phenomena…

计算与语言 · 计算机科学 2024-10-07 Yida Mu , Mali Jin , Xingyi Song , Nikolaos Aletras

Objective: This review aims to analyze the application of natural language processing (NLP) techniques in cancer research using electronic health records (EHRs) and clinical notes. This review addresses gaps in the existing literature by…

计算与语言 · 计算机科学 2025-02-05 Muhammad Bilal , Ameer Hamza , Nadia Malik

Large datasets have become commonplace in NLP research. However, the increased emphasis on data quantity has made it challenging to assess the quality of data. We introduce Data Maps---a model-based tool to characterize and diagnose…

计算与语言 · 计算机科学 2020-10-16 Swabha Swayamdipta , Roy Schwartz , Nicholas Lourie , Yizhong Wang , Hannaneh Hajishirzi , Noah A. Smith , Yejin Choi