English
Related papers

Related papers: Translating Hanja Historical Documents to Contempo…

200 papers

Historical records in Korea before the 20th century were primarily written in Hanja, an extinct language based on Chinese characters and not understood by modern Korean or Chinese speakers. Historians with expertise in this time period have…

Computation and Language · Computer Science 2022-10-12 Haneul Yoo , Jiho Jin , Juhee Son , JinYeong Bak , Kyunghyun Cho , Alice Oh

While Korean historical documents are invaluable cultural heritage, understanding those documents requires in-depth Hanja expertise. Hanja is an ancient language used in Korea before the 20th century, whose characters were borrowed from old…

Computation and Language · Computer Science 2025-01-22 Seyoung Song , Haneul Yoo , Jiho Jin , Kyunghyun Cho , Alice Oh

Historical documents in the Sinosphere are known to share common formats and practices, particularly in veritable records compiled by court historians. This shared linguistic heritage has led researchers to use Classical Chinese resources…

Computation and Language · Computer Science 2026-03-24 Seyoung Song , Haneul Yoo , Jiho Jin , Kyunghyun Cho , Alice Oh

Understanding voluminous historical records provides clues on the past in various aspects, such as social and political issues and even natural science facts. However, it is generally difficult to fully utilize the historical records, since…

Computation and Language · Computer Science 2021-05-10 Kyeongpil Kang , Kyohoon Jin , Soyoung Yang , Sujin Jang , Jaegul Choo , Youngbin Kim

The history of the Korean language is characterized by a discrepancy between its spoken and written forms and a pivotal shift from Chinese characters to the Hangul alphabet. However, this linguistic evolution has remained largely unexplored…

Computation and Language · Computer Science 2026-05-04 Seyoung Song , Nawon Kim , Songeun Chae , Kiwoong Park , Jiho Jin , Haneul Yoo , Kyunghyun Cho , Alice Oh

We implemented a high-performance optical character recognition model for classical handwritten documents using data augmentation with highly variable cropping within the document region. Optical character recognition in handwritten…

Computer Vision and Pattern Recognition · Computer Science 2024-12-17 Joonmo Ahna , Taehong Jang , Quan Fengnyu , Hyungil Lee , Jaehyuk Lee , Sojung Lucia Kim

A named entity recognition and classification plays the first and foremost important role in capturing semantics in data and anchoring in translation as well as downstream study for history. However, NER in historical text has faced…

Computation and Language · Computer Science 2023-06-27 Sojung Lucia Kim , Taehong Jang , Joonmo Ahn , Hyungil Lee , Jaehyuk Lee

Deep learning-based approaches for automatic document layout analysis and content extraction have the potential to unlock rich information trapped in historical documents on a large scale. One major hurdle is the lack of large datasets for…

Computer Vision and Pattern Recognition · Computer Science 2020-04-21 Zejiang Shen , Kaixuan Zhang , Melissa Dell

Recent studies in natural language processing (NLP) have focused on modern languages and achieved state-of-the-art results in many tasks. Meanwhile, little attention has been paid to ancient texts and related tasks. Classical Chinese first…

Computation and Language · Computer Science 2024-07-03 Hao Wang , Hirofumi Shimizu , Daisuke Kawahara

Japan is a unique country with a distinct cultural heritage, which is reflected in billions of historical documents that have been preserved. However, the change in Japanese writing system in 1900 made these documents inaccessible for the…

Computation and Language · Computer Science 2021-06-15 Alex Lamb , Tarin Clanuwat , Siyu Han , Mikel Bober-Irizar , Asanobu Kitamoto

We propose a simple yet effective approach for improving Korean word representations using additional linguistic annotation (i.e. Hanja). We employ cross-lingual transfer learning in training word representations by leveraging the fact that…

Computation and Language · Computer Science 2019-11-01 Kang Min Yoo , Taeuk Kim , Sang-goo Lee

Despite the extensive applications of relation extraction (RE) tasks in various domains, little has been explored in the historical context, which contains promising data across hundreds and thousands of years. To promote the historical RE…

Computation and Language · Computer Science 2023-07-11 Soyoung Yang , Minseok Choi , Youngwoo Cho , Jaegul Choo

Accessibility to historical documents is mostly limited to scholars. This is due to the language barrier inherent in human language and the linguistic properties of these documents. Given a historical document, modernization aims to…

Computation and Language · Computer Science 2020-03-05 Miguel Domingo , Francisco Casacuberta

The field of Natural Language Processing (NLP) has seen significant advancements with the development of Large Language Models (LLMs). However, much of this research remains focused on English, often overlooking low-resource languages like…

Computation and Language · Computer Science 2024-08-22 Anh-Dung Vo , Minseong Jung , Wonbeen Lee , Daewoo Choi

Large language models (LLMs) trained on massive corpora demonstrate impressive capabilities in a wide range of tasks. While there are ongoing efforts to adapt these models to languages beyond English, the attention given to their evaluation…

Computation and Language · Computer Science 2024-03-21 Guijin Son , Hanwool Lee , Suwan Kim , Huiseo Kim , Jaecheol Lee , Je Won Yeom , Jihyu Jung , Jung Woo Kim , Songseong Kim

Large language models (LLMs) often show poor performance in low-resource languages like Korean, partly due to unique linguistic challenges such as homophonous Sino-Korean words that are indistinguishable in Hangul script. To address this…

Computation and Language · Computer Science 2025-07-16 Seungho Choi

Document-level Relation Extraction (DocRE) is the task of extracting all semantic relationships from a document. While studies have been conducted on English DocRE, limited attention has been given to DocRE in non-English languages. This…

Computation and Language · Computer Science 2024-04-26 Youmi Ma , An Wang , Naoaki Okazaki

Translating knowledge-intensive and entity-rich text between English and Korean requires transcreation to preserve language-specific and cultural nuances beyond literal, phonetic or word-for-word conversion. We evaluate 13 models (LLMs and…

Computation and Language · Computer Science 2025-04-30 Daniel Lee , Harsh Sharma , Jieun Han , Sunny Jeong , Alice Oh , Vered Shwartz

We present our demonstration of two machine translation applications to historical documents. The first task consists in generating a new version of a historical document, written in the modern version of its original language. The second…

Computation and Language · Computer Science 2021-02-03 Miguel Domingo , Francisco Casacuberta

Jejueo was classified as critically endangered by UNESCO in 2010. Although diverse efforts to revitalize it have been made, there have been few computational approaches. Motivated by this, we construct two new Jejueo datasets: Jejueo…

Computation and Language · Computer Science 2019-11-28 Kyubyong Park , Yo Joong Choe , Jiyeon Ham
‹ Prev 1 2 3 10 Next ›