中文
相关论文

相关论文: Decolonising Data Systems: Using Jyutping or Pinyi…

200 篇论文

Disambiguating scholars with identical names is essential for accurate authorship assignment and robust large-scale scientometric research. Existing methods are often designed for Latin-script metadata and perform poorly on Chinese names.…

数字图书馆 · 计算机科学 2026-04-07 Mingrong She , Liuhuaying Yang , Ana Maria Jaramillo , Lisette Espín-Noboa

Chinese Spelling Correction (CSC) aims to detect and correct erroneous characters in Chinese texts. Although efforts have been made to introduce phonetic information (Hanyu Pinyin) in this task, they typically merge phonetic representations…

计算与语言 · 计算机科学 2023-05-25 Zihong Liang , Xiaojun Quan , Qifan Wang

Achieving gender equality is a pivotal factor in realizing the UN's Global Goals for Sustainable Development. Gender bias studies work towards this and rely on name-based gender inference tools to assign individual gender labels when gender…

计算与语言 · 计算机科学 2024-05-13 Xiaocong Du , Haipeng Zhang

While GPT has become the de-facto method for text generation tasks, its application to pinyin input method remains unexplored. In this work, we make the first exploration to leverage Chinese GPT for pinyin input method. We find that a…

计算与语言 · 计算机科学 2022-03-03 Minghuan Tan , Yong Dai , Duyu Tang , Zhangyin Feng , Guoping Huang , Jing Jiang , Jiwei Li , Shuming Shi

Recent studies have demonstrated the efficacy of large language models (LLMs) in error correction for automatic speech recognition (ASR). However, much of the research focuses on the English language. This paper redirects the attention to…

计算与语言 · 计算机科学 2024-09-25 Zhiyuan Tang , Dong Wang , Shen Huang , Shidong Shang

For many real-world applications, the user-generated inputs usually contain various noises due to speech recognition errors caused by linguistic variations1 or typographical errors (typos). Thus, it is crucial to test model performance on…

计算与语言 · 计算机科学 2023-05-26 Chenglei Si , Zhengyan Zhang , Yingfa Chen , Xiaozhi Wang , Zhiyuan Liu , Maosong Sun

One of the challenges with finetuning pretrained language models (PLMs) is that their tokenizer is optimized for the language(s) it was pretrained on, but brittle when it comes to previously unseen variations in the data. This can for…

计算与语言 · 计算机科学 2023-04-21 Verena Blaschke , Hinrich Schütze , Barbara Plank

Disease name normalization is an important task in the medical domain. It classifies disease names written in various formats into standardized names, serving as a fundamental component in smart healthcare systems for various…

计算与语言 · 计算机科学 2024-06-14 Wenqian Cui , Xiangling Fu , Shaohui Liu , Mingjun Gu , Xien Liu , Ji Wu , Irwin King

While the capabilities of Large Language Models (LLMs) have been studied in both Simplified and Traditional Chinese, it is yet unclear whether LLMs exhibit differential performance when prompted in these two variants of written Chinese.…

计算与语言 · 计算机科学 2025-05-29 Hanjia Lyu , Jiebo Luo , Jian Kang , Allison Koenecke

Conversion of Chinese Grapheme-to-Phoneme (G2P) plays an important role in Mandarin Chinese Text-To-Speech (TTS) systems, where one of the biggest challenges is the task of polyphone disambiguation. Most of the previous polyphone…

声音 · 计算机科学 2022-11-18 Chunyu Qiang , Peng Yang , Hao Che , Jinba Xiao , Xiaorui Wang , Zhongyuan Wang

When compiling databases, for example to meet the needs of healthcare establishments, there is quite a common problem with the introduction and further processing of names and last names of doctors and patients that are highly specialized…

计算与语言 · 计算机科学 2019-11-04 V. Buriachok , M. Hadzhyiev , V. Sokolov , P. Skladannyi , L. Kuzmenko

Disease name normalization is an important task in the medical domain. It classifies disease names written in various formats into standardized names, serving as a fundamental component in smart healthcare systems for various…

计算与语言 · 计算机科学 2025-01-03 Wenqian Cui , Xiangling Fu , Shaohui Liu , Mingjun Gu , Xien Liu , Ji Wu , Irwin King

Approximate string-matching methods to account for complex variation in highly discriminatory text fields, such as personal names, can enhance probabilistic record linkage. However, discriminating between matching and non-matching strings…

Conversion of Chinese graphemes to phonemes (G2P) is an essential component in Mandarin Chinese Text-To-Speech (TTS) systems. One of the biggest challenges in Chinese G2P conversion is how to disambiguate the pronunciation of polyphones -…

计算与语言 · 计算机科学 2020-09-18 Kyubyong Park , Seanie Lee

Simplified Chinese to Traditional Chinese character conversion is a common preprocessing step in Chinese NLP. Despite this, current approaches have poor performance because they do not take into account that a simplified Chinese character…

计算与语言 · 计算机科学 2020-05-08 Pranav A , Isabelle Augenstein

Privacy protection raises great attention on both legal levels and user awareness. To protect user privacy, countries enact laws and regulations requiring software privacy policies to regulate their behavior. However, privacy policies are…

密码学与安全 · 计算机科学 2022-12-09 Kaifa Zhao , Le Yu , Shiyao Zhou , Jing Li , Xiapu Luo , Yat Fei Aemon Chiu , Yutong Liu

The Chinese pronunciation system offers two characteristics that distinguish it from other languages: deep phonemic orthography and intonation variations. We are the first to argue that these two important properties can play a major role…

计算与语言 · 计算机科学 2019-01-24 Haiyun Peng , Yukun Ma , Soujanya Poria , Yang Li , Erik Cambria

Chinese word segmentation (CWS) is often regarded as a character-based sequence labeling task in most current works which have achieved great success with the help of powerful neural networks. However, these works neglect an important clue:…

计算与语言 · 计算机科学 2019-05-31 Jingkang Wang , Jianing Zhou , Jie Zhou , Gongshen Liu

Transliteration is a task of translating named entities from a language to another, based on phonetic similarity. The task has embraced deep learning approaches in recent years, yet, most ignore the phonetic features of the involved…

计算与语言 · 计算机科学 2022-01-24 Shi Cheng , Zhuofei Ding , Songpeng Yan

Phonetic Cloaking Replacement (PCR), defined as the deliberate use of homophonic or near-homophonic variants to hide toxic intent, has become a major obstacle to Chinese content moderation. While this problem is well-recognized, existing…

计算与语言 · 计算机科学 2025-07-11 Haotan Guo , Jianfei He , Jiayuan Ma , Hongbin Na , Zimu Wang , Haiyang Zhang , Qi Chen , Wei Wang , Zijing Shi , Tao Shen , Ling Chen
‹ 上一页 1 2 3 10 下一页 ›