English
Related papers

Related papers: A Topic-aware Comparable Corpus of Chinese Variati…

200 papers

This paper introduces a new open-sourced Mandarin speech corpus, called DiDiSpeech. It consists of about 800 hours of speech data at 48kHz sampling rate from 6000 speakers and the corresponding texts. All speech data in the corpus is…

Audio and Speech Processing · Electrical Eng. & Systems 2021-02-09 Tingwei Guo , Cheng Wen , Dongwei Jiang , Ne Luo , Ruixiong Zhang , Shuaijiang Zhao , Wubo Li , Cheng Gong , Wei Zou , Kun Han , Xiangang Li

Taiwanese China Studies (CS) has developed into a rich, interdisciplinary research field shaped by the unique geopolitical position and long standing academic engagement with Mainland China. This study responds to the growing need to…

Artificial Intelligence · Computer Science 2025-05-16 Hsuan-Lei Shao

Chinese features prominently in the Chinese communities located in the nations of Malay Archipelago. In these countries, Chinese has undergone the process of adjustment to the local languages and cultures, which leads to the occurrence of a…

Computation and Language · Computer Science 2022-09-13 Nankai Lin , Sihui Fu , Hongyan Wu , Shengyi Jiang

In this paper, we aim to address the challenges surrounding the translation of ancient Chinese text: (1) The linguistic gap due to the difference in eras results in translations that are poor in quality, and (2) most translations are…

Computation and Language · Computer Science 2021-07-08 Ernie Chang , Yow-Ting Shiue , Hui-Syuan Yeh , Vera Demberg

Adpositions are frequent markers of semantic relations, but they are highly ambiguous and vary significantly from language to language. Moreover, there is a dearth of annotated corpora for investigating the cross-linguistic variation of…

Digital Libraries · Computer Science 2020-03-20 Siyao Peng , Yang Liu , Yilun Zhu , Austin Blodgett , Yushi Zhao , Nathan Schneider

In this paper, we introduce the Chinese corpus from CLUE organization, CLUECorpus2020, a large-scale corpus that can be used directly for self-supervised learning such as pre-training of a language model, or language generation. It has 100G…

Computation and Language · Computer Science 2020-03-06 Liang Xu , Xuanwei Zhang , Qianqian Dong

This study adapts Semantic Network of Adposition and Case Supersenses (SNACS) annotation to Mandarin Chinese and demonstrates that the same supersense categories are appropriate for Chinese adposition semantics. We annotated 15 chapters of…

Computation and Language · Computer Science 2019-04-25 Yilun Zhu , Yang Liu , Siyao Peng , Austin Blodgett , Yushi Zhao , Nathan Schneider

Current large language models demonstrate deficiencies in understanding low-resource languages, particularly the minority languages in China. This limitation stems from the scarcity of available pre-training data. To address this…

Computation and Language · Computer Science 2024-06-14 Chen Zhang , Mingxu Tao , Quzhe Huang , Jiuheng Lin , Zhibin Chen , Yansong Feng

Topic modelling is a pivotal unsupervised machine learning technique for extracting valuable insights from large document collections. Existing neural topic modelling methods often encode contextual information of documents, while ignoring…

Computation and Language · Computer Science 2025-02-07 Yanan Ma , Chenghao Xiao , Chenhan Yuan , Sabine N van der Veer , Lamiece Hassan , Chenghua Lin , Goran Nenadic

Ancient Chinese texts present an area of enormous challenge and opportunity for humanities scholars interested in exploiting computational methods to assist in the development of new insights and interpretations of culturally significant…

Computation and Language · Computer Science 2017-02-06 Colin Allen , Hongliang Luo , Jaimie Murdock , Jianghuai Pu , Xiaohong Wang , Yanjie Zhai , Kun Zhao

The emergence of social media has largely eased the way people receive information and participate in public discussions. However, in countries with strict regulations on discussions in the public space, social media is no exception. To…

Social and Information Networks · Computer Science 2022-07-26 Yichi Qian , Qiyi Shan , Hanjia Lyu , Jiebo Luo

Multilingual topic models enable crosslingual tasks by extracting consistent topics from multilingual corpora. Most models require parallel or comparable training corpora, which limits their ability to generalize. In this paper, we first…

Computation and Language · Computer Science 2018-06-13 Shudong Hao , Michael J. Paul

We describe a method of using statistically-collected Chinese character groups from a corpus to augment a Chinese dictionary. The method is particularly useful for extracting domain-specific and regional words not readily available in…

cmp-lg · Computer Science 2008-02-03 Pascale Fung , Dekai Wu

This technical report presents our initial attempt to build a spoken large language model (LLM) for Taiwanese Mandarin, specifically tailored to enable real-time, speech-to-speech interaction in multi-turn conversations. Our end-to-end…

Multilingual pre-trained language models have shown impressive performance on cross-lingual tasks. It greatly facilitates the applications of natural language processing on low-resource languages. However, there are still some languages…

Computation and Language · Computer Science 2022-09-22 Ziqing Yang , Zihang Xu , Yiming Cui , Baoxin Wang , Min Lin , Dayong Wu , Zhigang Chen

We present measurements and analysis of censorship on Weibo, a popular microblogging site in China. Since we were limited in the rate at which we could download posts, we identified users likely to participate in sensitive topics and…

Information Retrieval · Computer Science 2012-11-28 Tao Zhu , David Phipps , Adam Pridgen , Jedidiah R. Crandall , Dan S. Wallach

The evaluation of large language models (LLMs) has drawn substantial attention in the field recently. This work focuses on evaluating LLMs in a Chinese context, specifically, for Traditional Chinese which has been largely underrepresented…

Computation and Language · Computer Science 2024-04-01 Po-Heng Chen , Sijia Cheng , Wei-Lin Chen , Yen-Ting Lin , Yun-Nung Chen

Cued Speech (CS) is a communication system developed for deaf people, which exploits hand cues to complement speechreading at the phonetic level. Currently, it is estimated that CS has been adapted to over 60 languages; however, no official…

Audio and Speech Processing · Electrical Eng. & Systems 2020-01-06 Liu Li , Feng Gang

We present BreezyVoice, a Text-to-Speech (TTS) system specifically adapted for Taiwanese Mandarin, highlighting phonetic control abilities to address the unique challenges of polyphone disambiguation in the language. Building upon…

This paper proposes a modularized sense induction and representation learning model that jointly learns bilingual sense embeddings that align well in the vector space, where the cross-lingual signal in the English-Chinese parallel corpus is…

Computation and Language · Computer Science 2018-10-23 Ta-Chung Chi , Yun-Nung Chen