中文
相关论文

相关论文: A Topic-aware Comparable Corpus of Chinese Variati…

200 篇论文

This paper introduces a new open-sourced Mandarin speech corpus, called DiDiSpeech. It consists of about 800 hours of speech data at 48kHz sampling rate from 6000 speakers and the corresponding texts. All speech data in the corpus is…

音频与语音处理 · 电气工程与系统科学 2021-02-09 Tingwei Guo , Cheng Wen , Dongwei Jiang , Ne Luo , Ruixiong Zhang , Shuaijiang Zhao , Wubo Li , Cheng Gong , Wei Zou , Kun Han , Xiangang Li

Taiwanese China Studies (CS) has developed into a rich, interdisciplinary research field shaped by the unique geopolitical position and long standing academic engagement with Mainland China. This study responds to the growing need to…

人工智能 · 计算机科学 2025-05-16 Hsuan-Lei Shao

Chinese features prominently in the Chinese communities located in the nations of Malay Archipelago. In these countries, Chinese has undergone the process of adjustment to the local languages and cultures, which leads to the occurrence of a…

计算与语言 · 计算机科学 2022-09-13 Nankai Lin , Sihui Fu , Hongyan Wu , Shengyi Jiang

In this paper, we aim to address the challenges surrounding the translation of ancient Chinese text: (1) The linguistic gap due to the difference in eras results in translations that are poor in quality, and (2) most translations are…

计算与语言 · 计算机科学 2021-07-08 Ernie Chang , Yow-Ting Shiue , Hui-Syuan Yeh , Vera Demberg

Adpositions are frequent markers of semantic relations, but they are highly ambiguous and vary significantly from language to language. Moreover, there is a dearth of annotated corpora for investigating the cross-linguistic variation of…

数字图书馆 · 计算机科学 2020-03-20 Siyao Peng , Yang Liu , Yilun Zhu , Austin Blodgett , Yushi Zhao , Nathan Schneider

In this paper, we introduce the Chinese corpus from CLUE organization, CLUECorpus2020, a large-scale corpus that can be used directly for self-supervised learning such as pre-training of a language model, or language generation. It has 100G…

计算与语言 · 计算机科学 2020-03-06 Liang Xu , Xuanwei Zhang , Qianqian Dong

This study adapts Semantic Network of Adposition and Case Supersenses (SNACS) annotation to Mandarin Chinese and demonstrates that the same supersense categories are appropriate for Chinese adposition semantics. We annotated 15 chapters of…

计算与语言 · 计算机科学 2019-04-25 Yilun Zhu , Yang Liu , Siyao Peng , Austin Blodgett , Yushi Zhao , Nathan Schneider

Current large language models demonstrate deficiencies in understanding low-resource languages, particularly the minority languages in China. This limitation stems from the scarcity of available pre-training data. To address this…

计算与语言 · 计算机科学 2024-06-14 Chen Zhang , Mingxu Tao , Quzhe Huang , Jiuheng Lin , Zhibin Chen , Yansong Feng

Topic modelling is a pivotal unsupervised machine learning technique for extracting valuable insights from large document collections. Existing neural topic modelling methods often encode contextual information of documents, while ignoring…

计算与语言 · 计算机科学 2025-02-07 Yanan Ma , Chenghao Xiao , Chenhan Yuan , Sabine N van der Veer , Lamiece Hassan , Chenghua Lin , Goran Nenadic

Ancient Chinese texts present an area of enormous challenge and opportunity for humanities scholars interested in exploiting computational methods to assist in the development of new insights and interpretations of culturally significant…

计算与语言 · 计算机科学 2017-02-06 Colin Allen , Hongliang Luo , Jaimie Murdock , Jianghuai Pu , Xiaohong Wang , Yanjie Zhai , Kun Zhao

The emergence of social media has largely eased the way people receive information and participate in public discussions. However, in countries with strict regulations on discussions in the public space, social media is no exception. To…

社会与信息网络 · 计算机科学 2022-07-26 Yichi Qian , Qiyi Shan , Hanjia Lyu , Jiebo Luo

Multilingual topic models enable crosslingual tasks by extracting consistent topics from multilingual corpora. Most models require parallel or comparable training corpora, which limits their ability to generalize. In this paper, we first…

计算与语言 · 计算机科学 2018-06-13 Shudong Hao , Michael J. Paul

We describe a method of using statistically-collected Chinese character groups from a corpus to augment a Chinese dictionary. The method is particularly useful for extracting domain-specific and regional words not readily available in…

cmp-lg · 计算机科学 2008-02-03 Pascale Fung , Dekai Wu

This technical report presents our initial attempt to build a spoken large language model (LLM) for Taiwanese Mandarin, specifically tailored to enable real-time, speech-to-speech interaction in multi-turn conversations. Our end-to-end…

Multilingual pre-trained language models have shown impressive performance on cross-lingual tasks. It greatly facilitates the applications of natural language processing on low-resource languages. However, there are still some languages…

计算与语言 · 计算机科学 2022-09-22 Ziqing Yang , Zihang Xu , Yiming Cui , Baoxin Wang , Min Lin , Dayong Wu , Zhigang Chen

We present measurements and analysis of censorship on Weibo, a popular microblogging site in China. Since we were limited in the rate at which we could download posts, we identified users likely to participate in sensitive topics and…

信息检索 · 计算机科学 2012-11-28 Tao Zhu , David Phipps , Adam Pridgen , Jedidiah R. Crandall , Dan S. Wallach

The evaluation of large language models (LLMs) has drawn substantial attention in the field recently. This work focuses on evaluating LLMs in a Chinese context, specifically, for Traditional Chinese which has been largely underrepresented…

计算与语言 · 计算机科学 2024-04-01 Po-Heng Chen , Sijia Cheng , Wei-Lin Chen , Yen-Ting Lin , Yun-Nung Chen

Cued Speech (CS) is a communication system developed for deaf people, which exploits hand cues to complement speechreading at the phonetic level. Currently, it is estimated that CS has been adapted to over 60 languages; however, no official…

音频与语音处理 · 电气工程与系统科学 2020-01-06 Liu Li , Feng Gang

We present BreezyVoice, a Text-to-Speech (TTS) system specifically adapted for Taiwanese Mandarin, highlighting phonetic control abilities to address the unique challenges of polyphone disambiguation in the language. Building upon…

This paper proposes a modularized sense induction and representation learning model that jointly learns bilingual sense embeddings that align well in the vector space, where the cross-lingual signal in the English-Chinese parallel corpus is…

计算与语言 · 计算机科学 2018-10-23 Ta-Chung Chi , Yun-Nung Chen