English
Related papers

Related papers: Multi-Modal Multi-Granularity Tokenizer for Chu Ba…

200 papers

Pre-trained language models such as BERT have exhibited remarkable performances in many tasks in natural language understanding (NLU). The tokens in the models are usually fine-grained in the sense that for languages like English they are…

Computation and Language · Computer Science 2021-05-28 Xinsong Zhang , Pengshuai Li , Hang Li

The choice of modeling units is crucial for automatic speech recognition (ASR) tasks. In mandarin scenarios, the Chinese characters represent meaning but are not directly related to the pronunciation. Thus only considering the writing of…

Computation and Language · Computer Science 2022-10-19 Yuting Yang , Binbin Du , Yuke Li

Chinese paleography, the study of ancient Chinese writing, is undergoing a computational turn powered by artificial intelligence. This position paper charts the trajectory of this emerging field, arguing that it is evolving from automating…

Computation and Language · Computer Science 2026-01-30 Yiran Rex Ma

Chinese pre-trained language models usually process text as a sequence of characters, while ignoring more coarse granularity, e.g., words. In this work, we propose a novel pre-training paradigm for Chinese -- Lattice-BERT, which explicitly…

Computation and Language · Computer Science 2021-05-31 Yuxuan Lai , Yijia Liu , Yansong Feng , Songfang Huang , Dongyan Zhao

The rapid growth of visual tokens in multimodal large language models (MLLMs) leads to excessive memory consumption and inference latency, especially when handling high-resolution images and videos. Token pruning is a technique used to…

Computer Vision and Pattern Recognition · Computer Science 2025-12-02 Zhongyu Yang , Dannong Xu , Wei Pang , Yingfang Yuan

Most existing online writer-identification systems require that the text content is supplied in advance and rely on separately designed features and classifiers. The identifications are based on lines of text, entire paragraphs, or entire…

Computer Vision and Pattern Recognition · Computer Science 2015-05-20 Weixin Yang , Lianwen Jin , Manfei Liu

The archaeological dating of bronze dings has played a critical role in the study of ancient Chinese history. Current archaeology depends on trained experts to carry out bronze dating, which is time-consuming and labor-intensive. For such…

Computer Vision and Pattern Recognition · Computer Science 2023-06-05 Rixin Zhou , Jiafu Wei , Qian Zhang , Ruihua Qi , Xi Yang , Chuntao Li

Bronze inscriptions (BI), engraved on ritual vessels, constitute a crucial stage of early Chinese writing and provide indispensable evidence for archaeological and historical studies. However, automatic BI recognition remains difficult due…

Computer Vision and Pattern Recognition · Computer Science 2025-10-03 Rixin Zhou , Peiqiang Qiu , Qian Zhang , Chuntao Li , Xi Yang

The information provided by historical documents has always been indispensable in the transmission of human civilization, but it has also made these books susceptible to damage due to various factors. Thanks to recent technology, the…

Computer Vision and Pattern Recognition · Computer Science 2021-04-06 Chia-Wei Tang , Chao-Lin Liu , Po-Sen Chiu

Recently, word enhancement has become very popular for Chinese Named Entity Recognition (NER), reducing segmentation errors and increasing the semantic and boundary information of Chinese words. However, these methods tend to ignore the…

Computation and Language · Computer Science 2021-07-13 Shuang Wu , Xiaoning Song , Zhenhua Feng

Pretrained language models (PLMs) have shown marvelous improvements across various NLP tasks. Most Chinese PLMs simply treat an input text as a sequence of characters, and completely ignore word information. Although Whole Word Masking can…

Computation and Language · Computer Science 2023-03-23 Xinnian Liang , Zefan Zhou , Hui Huang , Shuangzhi Wu , Tong Xiao , Muyun Yang , Zhoujun Li , Chao Bian

Chinese is one of the most widely used languages in the world, yet online handwritten Chinese character recognition (OLHCCR) remains challenging. To recognize Chinese characters, one popular choice is to adopt the 2D convolutional neural…

Computer Vision and Pattern Recognition · Computer Science 2020-04-21 Ji Gan , Weiqiang Wang , Ke Lu

Given the advantage and recent success of English character-level and subword-unit models in several NLP tasks, we consider the equivalent modeling problem for Chinese. Chinese script is logographic and many Chinese logograms are composed…

Computation and Language · Computer Science 2018-09-11 Falcon Z. Dai , Zheng Cai

We present TMMLU+, a new benchmark designed for Traditional Chinese language understanding. TMMLU+ is a multi-choice question-answering dataset with 66 subjects from elementary to professional level. It is six times larger and boasts a more…

Computation and Language · Computer Science 2024-07-12 Zhi-Rui Tam , Ya-Ting Pai , Yen-Wei Lee , Jun-Da Chen , Wei-Min Chu , Sega Cheng , Hong-Han Shuai

Identifying the different varieties of the same language is more challenging than unrelated languages identification. In this paper, we propose an approach to discriminate language varieties or dialects of Mandarin Chinese for the Mainland…

Computation and Language · Computer Science 2017-01-10 Fan Xu , Mingwen Wang , Maoxi Li

Ancient Chinese word segmentation (WSG) and part-of-speech tagging (POS) are important to study ancient Chinese, but the amount of ancient Chinese WSG and POS tagging data is still rare. In this paper, we propose a novel augmentation method…

Computation and Language · Computer Science 2023-03-07 Shuo Feng , Piji Li

Disambiguating scholars with identical names is essential for accurate authorship assignment and robust large-scale scientometric research. Existing methods are often designed for Latin-script metadata and perform poorly on Chinese names.…

Digital Libraries · Computer Science 2026-04-07 Mingrong She , Liuhuaying Yang , Ana Maria Jaramillo , Lisette Espín-Noboa

Oracle bone script, one of the earliest known forms of ancient Chinese writing, presents invaluable research materials for scholars studying the humanities and geography of the Shang Dynasty, dating back 3,000 years. The immense historical…

Computer Vision and Pattern Recognition · Computer Science 2024-09-04 Pengjie Wang , Kaile Zhang , Xinyu Wang , Shengwei Han , Yongge Liu , Jinpeng Wan , Haisu Guan , Zhebin Kuang , Lianwen Jin , Xiang Bai , Yuliang Liu

Multi-stroke characters in scripts such as Chinese and Japanese can be highly complex, posing significant challenges for both native speakers and, especially, non-native learners. If these characters can be simplified without degrading…

Computer Vision and Pattern Recognition · Computer Science 2025-07-01 Ryo Ishiyama , Shinnosuke Matsuo , Seiichi Uchida

Recent pretraining models in Chinese neglect two important aspects specific to the Chinese language: glyph and pinyin, which carry significant syntax and semantic information for language understanding. In this work, we propose ChineseBERT,…

Computation and Language · Computer Science 2021-07-01 Zijun Sun , Xiaoya Li , Xiaofei Sun , Yuxian Meng , Xiang Ao , Qing He , Fei Wu , Jiwei Li