English
Related papers

Related papers: Topic Modeling the H\`an di\u{a}n Ancient Classics

200 papers

This study aims to fill the gap by constructing a topic-aware comparable corpus of Mainland Chinese Mandarin and Taiwanese Mandarin from the social media in Mainland China and Taiwan, respectively. Using Dcard for Taiwanese Mandarin and…

Computation and Language · Computer Science 2024-11-19 Da-Chen Lian , Shu-Kai Hsieh

The scarcity of large-scale, open-source data for dialects severely hinders progress in speech technology, a challenge particularly acute for the widely spoken Sichuanese dialects of Chinese. To address this critical gap, we introduce…

Computation and Language · Computer Science 2025-09-23 Yuhang Dai , Ziyu Zhang , Shuai Wang , Longhao Li , Zhao Guo , Tianlun Zuo , Shuiyuan Wang , Hongfei Xue , Chengyou Wang , Qing Wang , Xin Xu , Hui Bu , Jie Li , Jian Kang , Binbin Zhang , Lei Xie

Scientific knowledge increasingly depends on complex computational processes where both hardware and software layers can influence research outcomes. As computational complexity grows, classical-quantum integration provides a lens for…

Emerging Technologies · Computer Science 2026-03-06 Anna Vrtiak , Duuk Baten , Ariana Torres-Knoop

The development of digital humanities has opened new perspectives in the history of Islam: whether dealing with thin or sometimes vast source corpora (such as S\=ira, al-Tbar\=i, al-Dahab\=i, etc.), these tools allow us to approach texts…

Digital Libraries · Computer Science 2024-04-25 Adrien de Jarmy

Chinese ancient documents, invaluable carriers of millennia of Chinese history and culture, hold rich knowledge across diverse fields but face challenges in digitization and understanding, i.e., traditional methods only scan images, while…

The data article presents the large bilingual parallel corpus of low-resourced language pair Sanskrit-Hindi, named SAHAAYAK 2023. The corpus contains total of 1.5M sentence pairs between Sanskrit and Hindi. To make the universal usability…

Computation and Language · Computer Science 2023-07-04 Vishvajitsinh Bakrola , Jitendra Nasariwala

This paper delves into the text processing aspects of Language Computing, which enables computers to understand, interpret, and generate human language. Focusing on tasks such as speech recognition, machine translation, sentiment analysis,…

Computation and Language · Computer Science 2024-08-13 Kengatharaiyer Sarveswaran

Fine-grained entity typing is a challenging task with wide applications. However, most existing datasets for this task are in English. In this paper, we introduce a corpus for Chinese fine-grained entity typing that contains 4,800 mentions…

Computation and Language · Computer Science 2020-04-21 Chin Lee , Hongliang Dai , Yangqiu Song , Xin Li

Historical visualizations are a valuable resource for studying the history of visualization and inspecting the cultural context where they were created. When investigating historical visualizations, it is essential to consider contributions…

Human-Computer Interaction · Computer Science 2025-02-27 Xiyao Mei , Yu Zhang , Chaofan Yang , Rui Shi , Xiaoru Yuan

Recent astonishing experiments with quantum computers have demonstrated unambiguously the existence of a quantum multiverse, where calculations of mind-boggling complexity are effortlessly computed in just a few minutes. Here, we…

Quantum Physics · Physics 2025-04-01 Brian R. La Cour , Noah A. Davis

With the rapid evolution of cross-strait situation, "Mainland China" as a subject of social science study has evoked the voice of "Rethinking China Study" among intelligentsia recently. This essay tried to apply an automatic content…

Digital Libraries · Computer Science 2023-06-22 Hsuan-Lei Shao , Sieh-Chuen Huang , Yun-Cheng Tsai

Recent studies in natural language processing (NLP) have focused on modern languages and achieved state-of-the-art results in many tasks. Meanwhile, little attention has been paid to ancient texts and related tasks. Classical Chinese first…

Computation and Language · Computer Science 2024-07-03 Hao Wang , Hirofumi Shimizu , Daisuke Kawahara

The birth and rapid development of large language models (LLMs) have caused quite a stir in the field of literature. Once considered unattainable, AI's role in literary creation is increasingly becoming a reality. In genres such as poetry,…

Computation and Language · Computer Science 2024-09-12 Cheng Zhao , Bin Wang , Zhen Wang

Tomb biographies of the Tang dynasty provide invaluable information about Chinese history. The original biographies are classical Chinese texts which contain neither word boundaries nor sentence boundaries. Relying on three published books…

Computation and Language · Computer Science 2019-08-29 Chao-Lin Liu , Yi Chang

The peer merit review of research proposals has been the major mechanism for deciding grant awards. However, research proposals have become increasingly interdisciplinary. It has been a longstanding challenge to assign interdisciplinary…

Information Retrieval · Computer Science 2023-02-23 Meng Xiao , Ziyue Qiao , Yanjie Fu , Hao Dong , Yi Du , Pengyang Wang , Hui Xiong , Yuanchun Zhou

Ancient scripts, e.g., Egyptian hieroglyphs, Oracle Bone Inscriptions, and Ancient Greek inscriptions, serve as vital carriers of human civilization, embedding invaluable historical and cultural information. Automating ancient script image…

Computer Vision and Pattern Recognition · Computer Science 2025-06-25 Xiaolei Diao , Rite Bo , Yanling Xiao , Lida Shi , Zhihan Zhou , Hao Xu , Chuntao Li , Xiongfeng Tang , Massimo Poesio , Cédric M. John , Daqian Shi

Human conversations are complicated and building a human-like dialogue agent is an extremely challenging task. With the rapid development of deep learning techniques, data-driven models become more and more prevalent which need a huge…

Computation and Language · Computer Science 2020-03-25 Meng Chen , Ruixue Liu , Lei Shen , Shaozu Yuan , Jingyan Zhou , Youzheng Wu , Xiaodong He , Bowen Zhou

Historical newspapers are a source of research for the human and social sciences. However, these image collections are difficult to read by machine due to the low quality of the print, the lack of standardization of the pages in addition to…

Information Retrieval · Computer Science 2020-02-21 José E. B. Maia , Gildácio J. de A. Sá

This paper proposes HistoLens, a multi-layered analysis framework for historical texts based on Large Language Models (LLMs). Using the important Western Han dynasty text "Yantie Lun" as a case study, we demonstrate the framework's…

Computation and Language · Computer Science 2024-11-18 Yifan Zeng

What are the best methods of capturing thematic similarity between literary texts? Knowing the answer to this question would be useful for automatic clustering of book genres, or any other thematic grouping. This paper compares a variety of…

Computation and Language · Computer Science 2023-05-22 Oleg Sobchuk , Artjoms Šeļa