中文
相关论文

相关论文: Shared Heritage, Distinct Writing: Rethinking Reso…

200 篇论文

Historical records in Korea before the 20th century were primarily written in Hanja, an extinct language based on Chinese characters and not understood by modern Korean or Chinese speakers. Historians with expertise in this time period have…

计算与语言 · 计算机科学 2022-10-12 Haneul Yoo , Jiho Jin , Juhee Son , JinYeong Bak , Kyunghyun Cho , Alice Oh

While Korean historical documents are invaluable cultural heritage, understanding those documents requires in-depth Hanja expertise. Hanja is an ancient language used in Korea before the 20th century, whose characters were borrowed from old…

计算与语言 · 计算机科学 2025-01-22 Seyoung Song , Haneul Yoo , Jiho Jin , Kyunghyun Cho , Alice Oh

We implemented a high-performance optical character recognition model for classical handwritten documents using data augmentation with highly variable cropping within the document region. Optical character recognition in handwritten…

计算机视觉与模式识别 · 计算机科学 2024-12-17 Joonmo Ahna , Taehong Jang , Quan Fengnyu , Hyungil Lee , Jaehyuk Lee , Sojung Lucia Kim

Recent studies in natural language processing (NLP) have focused on modern languages and achieved state-of-the-art results in many tasks. Meanwhile, little attention has been paid to ancient texts and related tasks. Classical Chinese first…

计算与语言 · 计算机科学 2024-07-03 Hao Wang , Hirofumi Shimizu , Daisuke Kawahara

The Annals of Joseon Dynasty (AJD) contain the daily records of the Kings of Joseon, the 500-year kingdom preceding the modern nation of Korea. The Annals were originally written in an archaic Korean writing system, `Hanja', and were…

计算与语言 · 计算机科学 2024-01-01 Juhee Son , Jiho Jin , Haneul Yoo , JinYeong Bak , Kyunghyun Cho , Alice Oh

The history of the Korean language is characterized by a discrepancy between its spoken and written forms and a pivotal shift from Chinese characters to the Hangul alphabet. However, this linguistic evolution has remained largely unexplored…

计算与语言 · 计算机科学 2026-05-04 Seyoung Song , Nawon Kim , Songeun Chae , Kiwoong Park , Jiho Jin , Haneul Yoo , Kyunghyun Cho , Alice Oh

We propose a simple yet effective approach for improving Korean word representations using additional linguistic annotation (i.e. Hanja). We employ cross-lingual transfer learning in training word representations by leveraging the fact that…

计算与语言 · 计算机科学 2019-11-01 Kang Min Yoo , Taeuk Kim , Sang-goo Lee

In Korean ancient documents, there is no spacing or punctuation, and they are written in classical Chinese characters. This makes it challenging for modern individuals and translation models to accurately interpret and translate them. While…

计算与语言 · 计算机科学 2023-12-20 Taehong Jang , Joonmo Ahn , Sojung Lucia Kim

In this paper, we aim to address the challenges surrounding the translation of ancient Chinese text: (1) The linguistic gap due to the difference in eras results in translations that are poor in quality, and (2) most translations are…

计算与语言 · 计算机科学 2021-07-08 Ernie Chang , Yow-Ting Shiue , Hui-Syuan Yeh , Vera Demberg

This study examines the cross-linguistic effectiveness of transfer learning for low-resource machine translation by fine-tuning models initially trained on typologically similar high-resource languages, using limited data from the target…

计算与语言 · 计算机科学 2025-09-03 Saughmon Boujkian

Large language models (LLMs) often show poor performance in low-resource languages like Korean, partly due to unique linguistic challenges such as homophonous Sino-Korean words that are indistinguishable in Hangul script. To address this…

计算与语言 · 计算机科学 2025-07-16 Seungho Choi

Korean-Chinese is a low resource language pair, but Korean and Chinese have a lot in common in terms of vocabulary. Sino-Korean words, which can be converted into corresponding Chinese characters, account for more than fifty of the entire…

计算与语言 · 计算机科学 2019-11-26 Jeonghyeok Park , Hai Zhao

Understanding voluminous historical records provides clues on the past in various aspects, such as social and political issues and even natural science facts. However, it is generally difficult to fully utilize the historical records, since…

计算与语言 · 计算机科学 2021-05-10 Kyeongpil Kang , Kyohoon Jin , Soyoung Yang , Sujin Jang , Jaegul Choo , Youngbin Kim

As pre-trained language models become more resource-demanding, the inequality between resource-rich languages such as English and resource-scarce languages is worsening. This can be attributed to the fact that the amount of available…

计算与语言 · 计算机科学 2022-09-15 Suhyune Son , Chanjun Park , Jungseob Lee , Midan Shim , Chanhee Lee , Yoonna Jang , Jaehyung Seo , Heuiseok Lim

Naively assuming English as a source language may hinder cross-lingual transfer for many languages by failing to consider the importance of language contact. Some languages are more well-connected than others, and target languages can…

计算与语言 · 计算机科学 2024-04-22 Hoang H. Nguyen , Chenwei Zhang , Ye Liu , Natalie Parde , Eugene Rohrbaugh , Philip S. Yu

Large language models (LLMs) trained on massive corpora demonstrate impressive capabilities in a wide range of tasks. While there are ongoing efforts to adapt these models to languages beyond English, the attention given to their evaluation…

计算与语言 · 计算机科学 2024-03-21 Guijin Son , Hanwool Lee , Suwan Kim , Huiseo Kim , Jaecheol Lee , Je Won Yeom , Jihyu Jung , Jung Woo Kim , Songseong Kim

Ancient people translated classical Chinese into Japanese using a system of annotations placed around characters. We abstract this process as sequence tagging tasks and fit them into modern language technologies. The research on this…

计算与语言 · 计算机科学 2026-01-22 Zilong Li , Jie Cao

Recent advances in Natural Language Processing (NLP) have underscored the crucial role of high-quality datasets in building large language models (LLMs). However, while extensive resources and analyses exist for English, the landscape for…

计算与语言 · 计算机科学 2025-10-16 Dasol Choi , Woomyoung Park , Youngsook Song

This study aims to compare three methods for translating ancient texts with sparse corpora: (1) the traditional statistical translation method of phrase alignment, (2) in-context LLM learning, and (3) proposed inter methodological approach…

计算与语言 · 计算机科学 2024-07-17 Sojung Lucia Kim , Taehong Jang , Joonmo Ahn

Detecting offensive language is a challenging task. Generalizing across different cultures and languages becomes even more challenging: besides lexical, syntactic and semantic differences, pragmatic aspects such as cultural norms and…

计算与语言 · 计算机科学 2023-04-03 Li Zhou , Laura Cabello , Yong Cao , Daniel Hershcovich
‹ 上一页 1 2 3 10 下一页 ›