English
Related papers

Related papers: HERITAGE: An End-to-End Web Platform for Processin…

200 papers

Realizing general-purpose language intelligence has been a longstanding goal for natural language processing, where standard evaluation benchmarks play a fundamental and guiding role. We argue that for general-purpose language intelligence…

Chinese ancient documents, invaluable carriers of millennia of Chinese history and culture, hold rich knowledge across diverse fields but face challenges in digitization and understanding, i.e., traditional methods only scan images, while…

We propose a Historical Document Reading Challenge on Large Chinese Structured Family Records, in short ICDAR2019 HDRC CHINESE. The objective of the proposed competition is to recognize and analyze the layout, and finally detect and…

Computer Vision and Pattern Recognition · Computer Science 2019-05-13 Rajkumar Saini , Derek Dobson , Jon Morrey , Marcus Liwicki , Foteini Simistira Liwicki

We introduce HRMCR (HAE-RAE Multi-Step Commonsense Reasoning), a benchmark designed to evaluate large language models' ability to perform multi-step reasoning in culturally specific contexts, focusing on Korean. The questions are…

Computation and Language · Computer Science 2025-03-13 Guijin Son , Hyunwoo Ko , Dasol Choi

While Large Language Models (LLMs) have increasingly assisted in historical tasks such as text processing, their capacity for professional-level historical reasoning remains underexplored. Existing benchmarks primarily assess basic…

Computation and Language · Computer Science 2026-04-28 Lirong Gao , Zeqing Wang , Yuyan Cai , Jiayi Deng , Yanmei Gu , Yiming Zhang , Jia Zhou , Yanfei Zhang , Junbo Zhao

South and North Korea both use the Korean language. However, Korean NLP research has focused on South Korean only, and existing NLP systems of the Korean language, such as neural machine translation (NMT) models, cannot properly handle…

Computation and Language · Computer Science 2022-01-28 Hwichan Kim , Sangwhan Moon , Naoaki Okazaki , Mamoru Komachi

Handwritten text recognition for historical documents is an important task but it remains difficult due to a lack of sufficient training data in combination with a large variability of writing styles and degradation of historical documents.…

Computer Vision and Pattern Recognition · Computer Science 2022-10-04 Christian M. Dahl , Torben S. D. Johansen , Emil N. Sørensen , Christian E. Westermann , Simon F. Wittrock

Cultural heritage is an arena of international relations that interests all states worldwide. The inscription process on the UNESCO World Heritage List and the UNESCO Representative List of the Intangible Cultural Heritage of Humanity often…

Taking advantage of computationally lightweight, but high-quality translators prompt consideration of new applications that address neglected languages. Locally run translators for less popular languages may assist data projects with…

Computation and Language · Computer Science 2021-01-15 David Noever , Josh Kalin , Matt Ciolino , Dom Hambrick , Gerry Dozier

This paper proposes HistoLens, a multi-layered analysis framework for historical texts based on Large Language Models (LLMs). Using the important Western Han dynasty text "Yantie Lun" as a case study, we demonstrate the framework's…

Computation and Language · Computer Science 2024-11-18 Yifan Zeng

Interpreting ancient Chinese has been the key to comprehending vast Chinese literature, tradition, and civilization. In this paper, we propose Erya for ancient Chinese translation. From a dataset perspective, we collect, clean, and classify…

Computation and Language · Computer Science 2023-08-02 Geyang Guo , Jiarong Yang , Fengyuan Lu , Jiaxin Qin , Tianyi Tang , Wayne Xin Zhao

English based datasets are commonly available from Kaggle, GitHub, or recently published papers. Although benchmark tests with English datasets are sufficient to show off the performances of new models and methods, still a researcher need…

Computation and Language · Computer Science 2022-11-29 Byunghyun Ban

Despite the significant success in the field of text recognition, complex and unsolved problems still exist in this field. In recent years, the recognition accuracy of the English language has greatly increased, while the problem of…

Computer Vision and Pattern Recognition · Computer Science 2019-12-04 Sergey A. Ilyuhin , Alexander V. Sheshkus , Vladimir L. Arlazarov

Large language models (LLMs) have showcased remarkable capabilities in understanding and generating language. However, their ability in comprehending ancient languages, particularly ancient Chinese, remains largely unexplored. To bridge…

Computation and Language · Computer Science 2023-10-17 Yixuan Zhang , Haonan Li

Chinese paper-cutting, an Intangible Cultural Heritage (ICH), faces challenges from the erosion of traditional culture due to the prevalence of realism alongside limited public access to cultural elements. While generative AI can enhance…

Human-Computer Interaction · Computer Science 2025-02-13 Huanchen Wang , Tianrun Qiu , Jiaping Li , Zhicong Lu , Yuxin Ma

Cross-language information retrieval (CLIR), where queries and documents are in different languages, needs a translation of queries and/or documents, so as to standardize both of them into a common representation. For this purpose, the use…

Computation and Language · Computer Science 2007-05-23 Atsushi Fujii , Tetsuya Ishikawa

The Manchu language, with its roots in the historical Manchurian region of Northeast China, is now facing a critical threat of extinction, as there are very few speakers left. In our efforts to safeguard the Manchu language, we introduce…

Computation and Language · Computer Science 2024-01-15 Jean Seo , Sungjoo Byun , Minha Kang , Sangah Lee

The Japanese writing system is complex, with three character types of Hiragana, Katakana, and Kanji. Kanji consists of thousands of unique characters, further adding to the complexity of character identification and literature…

Computer Vision and Pattern Recognition · Computer Science 2023-06-08 Angel I. Solis , Justin Zarkovacki , John Ly , Adham Atyabi

Sign Language Machine Translation (SLMT) aims to bridge communication between Deaf and hearing individuals. However, its progress is constrained by scarce datasets, limited signer diversity, and large domain gaps between sign motion…

Computation and Language · Computer Science 2026-03-23 Nada Shahin , Leila Ismail

Since 2006 we have undertaken to describe the differences between 17th century English and contemporary English thanks to NLP software. Studying a corpus spanning the whole century (tales of English travellers in the Ottoman Empire in the…

Computation and Language · Computer Science 2011-09-23 Odile Piton , Slim Mesfar , Hélène Pignot