中文
相关论文

相关论文: A Survey on Awesome Korean NLP Datasets

200 篇论文

Existing question answering systems mainly focus on dealing with text data. However, much of the data produced daily is stored in the form of tables that can be found in documents and relational databases, or on the web. To solve the task…

计算与语言 · 计算机科学 2022-05-03 Changwook Jun , Jooyoung Choi , Myoseop Sim , Hyun Kim , Hansol Jang , Kyungkoo Min

Instruction Tuning on Large Language Models is an essential process for model to function well and achieve high performance in specific tasks. Accordingly, in mainstream languages such as English, instruction-based datasets are being…

计算与语言 · 计算机科学 2024-03-26 Dongjun Jang , Sungjoo Byun , Hyemi Jo , Hyopil Shin

Natural language inference (NLI) and semantic textual similarity (STS) are key tasks in natural language understanding (NLU). Although several benchmark datasets for those tasks have been released in English and a few other languages, there…

计算与语言 · 计算机科学 2020-10-06 Jiyeon Ham , Yo Joong Choe , Kyubyong Park , Ilji Choi , Hyungjoon Soh

Although LLMs have made significant progress in various languages, there are still concerns about their effectiveness with low-resource agglutinative languages compared to languages such as English. In this study, we focused on Korean, a…

计算与语言 · 计算机科学 2025-07-08 Seunguk Yu , Kyeonghyun Kim , Jungmin Yun , Youngbin Kim

The scale, variety, and quantity of publicly-available NLP datasets has grown rapidly as researchers propose new tasks, larger models, and novel benchmarks. Datasets is a community library for contemporary NLP designed to support this…

A well-formulated benchmark plays a critical role in spurring advancements in the natural language processing (NLP) field, as it allows objective and precise evaluation of diverse models. As modern language models (LMs) have become more…

计算与语言 · 计算机科学 2022-04-12 Dohyeong Kim , Myeongjun Jang , Deuk Sin Kwon , Eric Davis

Recent advances in Natural Language Processing (NLP) have underscored the crucial role of high-quality datasets in building large language models (LLMs). However, while extensive resources and analyses exist for English, the landscape for…

计算与语言 · 计算机科学 2025-10-16 Dasol Choi , Woomyoung Park , Youngsook Song

South and North Korea both use the Korean language. However, Korean NLP research has focused on South Korean only, and existing NLP systems of the Korean language, such as neural machine translation (NMT) models, cannot properly handle…

计算与语言 · 计算机科学 2022-01-28 Hwichan Kim , Sangwhan Moon , Naoaki Okazaki , Mamoru Komachi

While the NLP community is generally aware of resource disparities among languages, we lack research that quantifies the extent and types of such disparity. Prior surveys estimating the availability of resources based on the number of…

计算与语言 · 计算机科学 2022-11-29 Xinyan Velocity Yu , Akari Asai , Trina Chatterjee , Junjie Hu , Eunsol Choi

Research in question answering datasets and models has gained a lot of attention in the research community. Many of them release their own question answering datasets as well as the models. There is tremendous progress that we have seen in…

计算与语言 · 计算机科学 2021-12-28 Andreas Chandra , Affandy Fahrizain , Ibrahim , Simon Willyanto Laufried

This paper introduces the Open Ko-LLM Leaderboard and the Ko-H5 Benchmark as vital tools for evaluating Large Language Models (LLMs) in Korean. Incorporating private test sets while mirroring the English Open LLM Leaderboard, we establish a…

计算与语言 · 计算机科学 2024-08-20 Chanjun Park , Hyeonwoo Kim , Dahyun Kim , Seonghwan Cho , Sanghoon Kim , Sukyung Lee , Yungi Kim , Hwalsuk Lee

Lyric translation, a field studied for over a century, is now attracting computational linguistics researchers. We identified two limitations in previous studies. Firstly, lyric translation studies have predominantly focused on Western…

计算与语言 · 计算机科学 2024-05-21 Haven Kim , Jongmin Jung , Dasaem Jeong , Juhan Nam

Large language models (LLMs) have demonstrated remarkable performance in the legal domain, with GPT-4 even passing the Uniform Bar Exam in the U.S. However their efficacy remains limited for non-standardized tasks and tasks in languages…

计算与语言 · 计算机科学 2024-10-14 Yeeun Kim , Young Rok Choi , Eunkyung Choi , Jinhwan Choi , Hai Jin Park , Wonseok Hwang

Despite the rapid development of large language models (LLMs) for the Korean language, there remains an obvious lack of benchmark datasets that test the requisite Korean cultural and linguistic knowledge. Because many existing Korean…

计算与语言 · 计算机科学 2024-07-08 Eunsu Kim , Juyoung Suk , Philhoon Oh , Haneul Yoo , James Thorne , Alice Oh

Scientific literature serves as a high-quality corpus, supporting a lot of Natural Language Processing (NLP) research. However, existing datasets are centered around the English language, which restricts the development of Chinese…

计算与语言 · 计算机科学 2022-09-13 Yudong Li , Yuqing Zhang , Zhe Zhao , Linlin Shen , Weijie Liu , Weiquan Mao , Hui Zhang

If someone is looking for a certain publication in the field of computer science, the searching person is likely to use the DBLP to find the desired publication. The DBLP data set is continuously extended with new publications, or rather…

计算与语言 · 计算机科学 2017-09-27 Paul Christian Sommerhoff

We propose KMMLU, a new Korean benchmark with 35,030 expert-level multiple-choice questions across 45 subjects ranging from humanities to STEM. While prior Korean benchmarks are translated from existing English benchmarks, KMMLU is…

The Open Ko-LLM Leaderboard has been instrumental in benchmarking Korean Large Language Models (LLMs), yet it has certain limitations. Notably, the disconnect between quantitative improvements on the overly academic leaderboard benchmarks…

计算与语言 · 计算机科学 2025-03-05 Hyeonwoo Kim , Dahyun Kim , Jihoo Kim , Sukyung Lee , Yungi Kim , Chanjun Park

Korean is often referred to as a low-resource language in the research community. While this claim is partially true, it is also because the availability of resources is inadequately advertised and curated. This work curates and reviews a…

计算与语言 · 计算机科学 2023-05-17 Won Ik Cho , Sangwhan Moon , Youngsook Song

This study constructed a Japanese chat dataset for tuning large language models (LLMs), which consist of about 8.4 million records. Recently, LLMs have been developed and gaining popularity. However, high-performing LLMs are usually mainly…

计算与语言 · 计算机科学 2023-05-23 Masanori Hirano , Masahiro Suzuki , Hiroki Sakaji
‹ 上一页 1 2 3 10 下一页 ›