English
Related papers

Related papers: HERITAGE: An End-to-End Web Platform for Processin…

200 papers

We introduce Korean Language Understanding Evaluation (KLUE) benchmark. KLUE is a collection of 8 Korean natural language understanding (NLU) tasks, including Topic Classification, SemanticTextual Similarity, Natural Language Inference,…

Large language models have exhibited significant enhancements in performance across various tasks. However, the complexity of their evaluation increases as these models generate more fluent and coherent content. Current multilingual…

Computation and Language · Computer Science 2024-12-11 Xiaonan Wang , Jinyoung Yeo , Joon-Ho Lim , Hansaem Kim

HERITRACE is a semantic data management system tailored for the GLAM sector. It is engineered to streamline data curation for non-technical users while also offering an efficient administrative interface for technical staff. The paper…

Digital Libraries · Computer Science 2024-04-25 Arcangelo Massari , Silvio Peroni

Despite the rapid development of large language models (LLMs) for the Korean language, there remains an obvious lack of benchmark datasets that test the requisite Korean cultural and linguistic knowledge. Because many existing Korean…

Computation and Language · Computer Science 2024-07-08 Eunsu Kim , Juyoung Suk , Philhoon Oh , Haneul Yoo , James Thorne , Alice Oh

The information provided by historical documents has always been indispensable in the transmission of human civilization, but it has also made these books susceptible to damage due to various factors. Thanks to recent technology, the…

Computer Vision and Pattern Recognition · Computer Science 2021-04-06 Chia-Wei Tang , Chao-Lin Liu , Po-Sen Chiu

HERITRACE is an open-source web application that enables users without Semantic Web expertise to curate RDF data through form-based interfaces with automatic provenance documentation and change tracking in RDF. It uses SHACL for data model…

Digital Libraries · Computer Science 2026-05-05 Arcangelo Massari , Silvio Peroni

In this paper, we propose an end-to-end trainable framework for restoring historical documents content that follows the correct reading order. In this framework, two branches named character branch and layout branch are added behind the…

Computer Vision and Pattern Recognition · Computer Science 2020-07-15 Weihong Ma , Hesuo Zhang , Lianwen Jin , Sihang Wu , Jiapeng Wang , Yongpan Wang

Cultural Heritage (CH) data hold invaluable knowledge, reflecting the history, traditions, and identities of societies, and shaping our understanding of the past and present. However, many CH collections contain outdated or offensive…

Computation and Language · Computer Science 2025-06-02 Orfeas Menis Mastromichalakis , Jason Liartis , Kristina Rose , Antoine Isaac , Giorgos Stamou

Japan is a unique country with a distinct cultural heritage, which is reflected in billions of historical documents that have been preserved. However, the change in Japanese writing system in 1900 made these documents inaccessible for the…

Computation and Language · Computer Science 2021-06-15 Alex Lamb , Tarin Clanuwat , Siyu Han , Mikel Bober-Irizar , Asanobu Kitamoto

This paper discusses the potential for integrating Generative Artificial Intelligence (GenAI) into professional heritage practice with the aim of enhancing the accessibility of public-facing guidance documents. We developed HAZEL, a GenAI…

Human-Computer Interaction · Computer Science 2025-10-17 Jessica Witte , Edmund Lee , Lisa Brausem , Verity Shillabeer , Chiara Bonacchi

Natural Language Processing (NLP) plays a pivotal role in the realm of Digital Humanities (DH) and serves as the cornerstone for advancing the structural analysis of historical and cultural heritage texts. This is particularly true for the…

Computation and Language · Computer Science 2024-04-23 Xuemei Tang , Zekun Deng , Qi Su , Hao Yang , Jun Wang

The field of Natural Language Processing (NLP) has seen significant advancements with the development of Large Language Models (LLMs). However, much of this research remains focused on English, often overlooking low-resource languages like…

Computation and Language · Computer Science 2024-08-22 Anh-Dung Vo , Minseong Jung , Wonbeen Lee , Daewoo Choi

Document-level Relation Extraction (DocRE) is the task of extracting all semantic relationships from a document. While studies have been conducted on English DocRE, limited attention has been given to DocRE in non-English languages. This…

Computation and Language · Computer Science 2024-04-26 Youmi Ma , An Wang , Naoaki Okazaki

Recognizing the full-page of Japanese historical documents is a challenging problem due to the complex layout/background and difficulty of writing styles, such as cursive and connected characters. Most of the previous methods divided the…

Computer Vision and Pattern Recognition · Computer Science 2020-05-07 Anh Duc Le

Classical Chinese, as the core carrier of Chinese culture, plays a crucial role in the inheritance and study of ancient literature. However, existing natural language processing models primarily optimize for Modern Chinese, resulting in…

Computation and Language · Computer Science 2025-04-30 Xinyu Yao , Mengdi Wang , Bo Chen , Xiaobing Zhao

Purpose: This study investigates the transcription principles underlying Hu\`it\'onggu\v{a}nx\`i Hu\'ay\'iy\`iy\v{u} (HHY), a series of multilingual glossaries compiled by the Ming government between the fifteenth and sixteenth centuries…

Computation and Language · Computer Science 2026-05-27 Ji-eun Kim

HERITRACE is a semantic data editor designed for cultural heritage institutions, addressing the gap between complex Semantic Web technologies and domain expert needs. ParaText Bibliographical Database, a specialized bibliographical database…

Digital Libraries · Computer Science 2025-08-22 Francesca Filograsso , Arcangelo Massari , Camillo Neri , Silvio Peroni

This memoir explores two fundamental aspects of Natural Language Processing (NLP): the creation of linguistic resources and the evaluation of NLP system performance. Over the past decade, my work has focused on developing a morpheme-based…

Computation and Language · Computer Science 2026-02-16 Jungyeul Park

Jejueo was classified as critically endangered by UNESCO in 2010. Although diverse efforts to revitalize it have been made, there have been few computational approaches. Motivated by this, we construct two new Jejueo datasets: Jejueo…

Computation and Language · Computer Science 2019-11-28 Kyubyong Park , Yo Joong Choe , Jiyeon Ham

The Korean writing system, \textit{Hangeul}, has a unique character representation rigidly following the invention principles recorded in \textit{Hunminjeongeum}.\footnote{\textit{Hunminjeongeum} is a book published in 1446 that describes…

Computation and Language · Computer Science 2026-04-28 SungHo Kim , Juhyeong Park , Yeachan Kim , SangKeun Lee