中文
相关论文

相关论文: HUE: Pretrained Model and Dataset for Understandin…

200 篇论文

While Korean historical documents are invaluable cultural heritage, understanding those documents requires in-depth Hanja expertise. Hanja is an ancient language used in Korea before the 20th century, whose characters were borrowed from old…

计算与语言 · 计算机科学 2025-01-22 Seyoung Song , Haneul Yoo , Jiho Jin , Kyunghyun Cho , Alice Oh

The Annals of Joseon Dynasty (AJD) contain the daily records of the Kings of Joseon, the 500-year kingdom preceding the modern nation of Korea. The Annals were originally written in an archaic Korean writing system, `Hanja', and were…

计算与语言 · 计算机科学 2024-01-01 Juhee Son , Jiho Jin , Haneul Yoo , JinYeong Bak , Kyunghyun Cho , Alice Oh

Historical documents in the Sinosphere are known to share common formats and practices, particularly in veritable records compiled by court historians. This shared linguistic heritage has led researchers to use Classical Chinese resources…

计算与语言 · 计算机科学 2026-03-24 Seyoung Song , Haneul Yoo , Jiho Jin , Kyunghyun Cho , Alice Oh

We implemented a high-performance optical character recognition model for classical handwritten documents using data augmentation with highly variable cropping within the document region. Optical character recognition in handwritten…

计算机视觉与模式识别 · 计算机科学 2024-12-17 Joonmo Ahna , Taehong Jang , Quan Fengnyu , Hyungil Lee , Jaehyuk Lee , Sojung Lucia Kim

Large language models (LLMs) trained on massive corpora demonstrate impressive capabilities in a wide range of tasks. While there are ongoing efforts to adapt these models to languages beyond English, the attention given to their evaluation…

计算与语言 · 计算机科学 2024-03-21 Guijin Son , Hanwool Lee , Suwan Kim , Huiseo Kim , Jaecheol Lee , Je Won Yeom , Jihyu Jung , Jung Woo Kim , Songseong Kim

The history of the Korean language is characterized by a discrepancy between its spoken and written forms and a pivotal shift from Chinese characters to the Hangul alphabet. However, this linguistic evolution has remained largely unexplored…

计算与语言 · 计算机科学 2026-05-04 Seyoung Song , Nawon Kim , Songeun Chae , Kiwoong Park , Jiho Jin , Haneul Yoo , Kyunghyun Cho , Alice Oh

A named entity recognition and classification plays the first and foremost important role in capturing semantics in data and anchoring in translation as well as downstream study for history. However, NER in historical text has faced…

计算与语言 · 计算机科学 2023-06-27 Sojung Lucia Kim , Taehong Jang , Joonmo Ahn , Hyungil Lee , Jaehyuk Lee

Recent advancements in Korean large language models (LLMs) have driven numerous benchmarks and evaluation methods, yet inconsistent protocols cause up to 10 p.p performance gaps across institutions. Overcoming these reproducibility gaps…

计算工程、金融与科学 · 计算机科学 2026-02-16 Hanwool Lee , Dasol Choi , Sooyong Kim , Ilgyun Jeong , Sangwon Baek , Guijin Son , Inseon Hwang , Naeun Lee , Seunghyeok Hong

Despite the extensive applications of relation extraction (RE) tasks in various domains, little has been explored in the historical context, which contains promising data across hundreds and thousands of years. To promote the historical RE…

计算与语言 · 计算机科学 2023-07-11 Soyoung Yang , Minseok Choi , Youngwoo Cho , Jaegul Choo

Understanding voluminous historical records provides clues on the past in various aspects, such as social and political issues and even natural science facts. However, it is generally difficult to fully utilize the historical records, since…

计算与语言 · 计算机科学 2021-05-10 Kyeongpil Kang , Kyohoon Jin , Soyoung Yang , Sujin Jang , Jaegul Choo , Youngbin Kim

We propose a simple yet effective approach for improving Korean word representations using additional linguistic annotation (i.e. Hanja). We employ cross-lingual transfer learning in training word representations by leveraging the fact that…

计算与语言 · 计算机科学 2019-11-01 Kang Min Yoo , Taeuk Kim , Sang-goo Lee

With the advent of Transformer, which was used in translation models in 2017, attention-based architectures began to attract attention. Furthermore, after the emergence of BERT, which strengthened the NLU-specific encoder part, which is a…

计算与语言 · 计算机科学 2021-12-07 Kichang Yang

Deep learning-based approaches for automatic document layout analysis and content extraction have the potential to unlock rich information trapped in historical documents on a large scale. One major hurdle is the lack of large datasets for…

计算机视觉与模式识别 · 计算机科学 2020-04-21 Zejiang Shen , Kaixuan Zhang , Melissa Dell

Recent advances in Natural Language Processing (NLP) have underscored the crucial role of high-quality datasets in building large language models (LLMs). However, while extensive resources and analyses exist for English, the landscape for…

计算与语言 · 计算机科学 2025-10-16 Dasol Choi , Woomyoung Park , Youngsook Song

Large language models (LLMs) often show poor performance in low-resource languages like Korean, partly due to unique linguistic challenges such as homophonous Sino-Korean words that are indistinguishable in Hangul script. To address this…

计算与语言 · 计算机科学 2025-07-16 Seungho Choi

South and North Korea both use the Korean language. However, Korean NLP research has focused on South Korean only, and existing NLP systems of the Korean language, such as neural machine translation (NMT) models, cannot properly handle…

计算与语言 · 计算机科学 2022-01-28 Hwichan Kim , Sangwhan Moon , Naoaki Okazaki , Mamoru Komachi

We introduce Korean Language Understanding Evaluation (KLUE) benchmark. KLUE is a collection of 8 Korean natural language understanding (NLU) tasks, including Topic Classification, SemanticTextual Similarity, Natural Language Inference,…

In Korean ancient documents, there is no spacing or punctuation, and they are written in classical Chinese characters. This makes it challenging for modern individuals and translation models to accurately interpret and translate them. While…

计算与语言 · 计算机科学 2023-12-20 Taehong Jang , Joonmo Ahn , Sojung Lucia Kim

Jejueo was classified as critically endangered by UNESCO in 2010. Although diverse efforts to revitalize it have been made, there have been few computational approaches. Motivated by this, we construct two new Jejueo datasets: Jejueo…

计算与语言 · 计算机科学 2019-11-28 Kyubyong Park , Yo Joong Choe , Jiyeon Ham

Instruction Tuning on Large Language Models is an essential process for model to function well and achieve high performance in specific tasks. Accordingly, in mainstream languages such as English, instruction-based datasets are being…

计算与语言 · 计算机科学 2024-03-26 Dongjun Jang , Sungjoo Byun , Hyemi Jo , Hyopil Shin
‹ 上一页 1 2 3 10 下一页 ›