中文
相关论文

相关论文: Multi-Modal Multi-Granularity Tokenizer for Chu Ba…

200 篇论文

Character-level Neural Machine Translation (NMT) models have recently achieved impressive results on many language pairs. They mainly do well for Indo-European language pairs, where the languages share the same writing system. However, for…

计算与语言 · 计算机科学 2018-08-28 Nikola I. Nikolov , Yuhuang Hu , Mi Xue Tan , Richard H. R. Hahnloser

Recently, it is quite common to integrate Chinese sequence labeling results to enhance syntactic and semantic parsing. However, little attention has been paid to the utility of hierarchy and structure information encoded in syntactic and…

计算与语言 · 计算机科学 2023-06-06 Xuemei Tang , Jun Wang , Qi Su

We collect nine corpora of representative Chinese poetry for the time span of 1046 BCE and 1644 CE for studying the history of Chinese words, collocations, and patterns. By flexibly integrating our own tools, we are able to provide new…

计算与语言 · 计算机科学 2017-09-19 Chao-Lin Liu

Although online handwriting verification has made great progress recently, the verification performances are still far behind the real usage owing to the small scale of the datasets as well as the limited biometric mediums. Therefore, this…

计算机视觉与模式识别 · 计算机科学 2022-11-28 Peirong Zhang , Jiajia Jiang , Yuliang Liu , Lianwen Jin

Comprehension of ancient texts plays an important role in archaeology and understanding of Chinese history and civilization. The rapid development of large language models needs benchmarks that can evaluate their comprehension of ancient…

计算与语言 · 计算机科学 2025-12-22 Zhihan Zhou , Daqian Shi , Rui Song , Lida Shi , Xiaolei Diao , Hao Xu

Different linguistic perspectives causes many diverse segmentation criteria for Chinese word segmentation (CWS). Most existing methods focus on improve the performance for each single criterion. However, it is interesting to exploit these…

计算与语言 · 计算机科学 2017-04-26 Xinchi Chen , Zhan Shi , Xipeng Qiu , Xuanjing Huang

In this paper, we aim to address the challenges surrounding the translation of ancient Chinese text: (1) The linguistic gap due to the difference in eras results in translations that are poor in quality, and (2) most translations are…

计算与语言 · 计算机科学 2021-07-08 Ernie Chang , Yow-Ting Shiue , Hui-Syuan Yeh , Vera Demberg

In social media, neural network models have been applied to hate speech detection, sentiment analysis, etc., but neural network models are susceptible to adversarial attacks. For instance, in a text classification task, the attacker…

计算与语言 · 计算机科学 2024-12-04 Xi Cao , Nuo Qun , Quzong Gesang , Yulei Zhu , Trashi Nyima

The Indus script is one of the major undeciphered scripts of the ancient world. The small size of the corpus, the absence of bilingual texts, and the lack of definite knowledge of the underlying language has frustrated efforts at…

计算与语言 · 计算机科学 2015-05-13 Nisha Yadav , Hrishikesh Joglekar , Rajesh P. N. Rao , M. N. Vahia , Iravatham Mahadevan , R. Adhikari

Chinese Spelling Correction (CSC) aims to detect and correct erroneous characters in Chinese texts. Although efforts have been made to introduce phonetic information (Hanyu Pinyin) in this task, they typically merge phonetic representations…

计算与语言 · 计算机科学 2023-05-25 Zihong Liang , Xiaojun Quan , Qifan Wang

Scene text recognition (STR) has been widely studied in academia and industry. Training a text recognition model often requires a large amount of labeled data, but data labeling can be difficult, expensive, or time-consuming, especially for…

计算机视觉与模式识别 · 计算机科学 2022-08-09 Yi-Chang Chen , Yu-Chuan Chang , Yen-Cheng Chang , Yi-Ren Yeh

Automatic analysis for modern Chinese has greatly improved the accuracy of text mining in related fields, but the study of ancient Chinese is still relatively rare. Ancient text division and lexical annotation are important parts of…

计算与语言 · 计算机科学 2023-10-13 Pengyu Wang , Zhichen Ren

Recently, language representation techniques have achieved great performances in text classification. However, most existing representation models are specifically designed for English materials, which may fail in Chinese because of the…

计算与语言 · 计算机科学 2022-12-19 Xunzhu Tang , Rujie Zhu , Tiezhu Sun , Shi Wang

Chinese character decomposition has been used as a feature to enhance Machine Translation (MT) models, combining radicals into character and word level models. Recent work has investigated ideograph or stroke level embedding. However,…

计算与语言 · 计算机科学 2021-04-12 Lifeng Han , Gareth J. F. Jones , Alan F. Smeaton , Paolo Bolzoni

A sequence-to-sequence learning with neural networks has empirically proven to be an effective framework for Chinese Spelling Correction (CSC), which takes a sentence with some spelling errors as input and outputs the corrected one.…

计算与语言 · 计算机科学 2021-06-02 Chong Li , Cenyuan Zhang , Xiaoqing Zheng , Xuanjing Huang

Constructing historical language models (LMs) plays a crucial role in aiding archaeological provenance studies and understanding ancient cultures. However, existing resources present major challenges for training effective LMs on historical…

计算与语言 · 计算机科学 2025-08-25 Xiaolei Diao , Zhihan Zhou , Lida Shi , Ting Wang , Ruihua Qi , Hao Xu , Daqian Shi

Oracle bone script (OBS), as China's earliest mature writing system, present significant challenges in automatic recognition due to their complex pictographic structures and divergence from modern Chinese characters. We introduce…

计算机视觉与模式识别 · 计算机科学 2024-11-28 Hanqi Jiang , Yi Pan , Junhao Chen , Zhengliang Liu , Yifan Zhou , Peng Shu , Yiwei Li , Huaqin Zhao , Stephen Mihm , Lewis C Howe , Tianming Liu

Tokenization is used almost universally by modern language models, enabling efficient text representation using multi-byte or multi-character tokens. However, prior work has shown that tokenization can introduce distortion into the model's…

计算与语言 · 计算机科学 2026-05-08 Jonathan Hayase , Alisa Liu , Noah A. Smith , Sewoong Oh

In this paper, we present an effective method to analyze the recognition confidence of handwritten Chinese character, based on the softmax regression score of a high performance convolutional neural networks (CNN). Through careful and…

计算机视觉与模式识别 · 计算机科学 2015-05-26 Meijun He , Shuye Zhang , Huiyun Mao , Lianwen Jin

In this paper, we develop a low than character feature embedding called radical embedding, and apply it on LSTM model for sentence segmentation of pre modern Chinese texts. The datasets includes over 150 classical Chinese books from 3…

计算与语言 · 计算机科学 2020-02-20 Xu Han , Hongsu Wang , Sanqian Zhang , Qunchao Fu , Jun S. Liu