中文
相关论文

相关论文: Multi-Modal Multi-Granularity Tokenizer for Chu Ba…

200 篇论文

Pre-trained language models such as BERT have exhibited remarkable performances in many tasks in natural language understanding (NLU). The tokens in the models are usually fine-grained in the sense that for languages like English they are…

计算与语言 · 计算机科学 2021-05-28 Xinsong Zhang , Pengshuai Li , Hang Li

The choice of modeling units is crucial for automatic speech recognition (ASR) tasks. In mandarin scenarios, the Chinese characters represent meaning but are not directly related to the pronunciation. Thus only considering the writing of…

计算与语言 · 计算机科学 2022-10-19 Yuting Yang , Binbin Du , Yuke Li

Chinese paleography, the study of ancient Chinese writing, is undergoing a computational turn powered by artificial intelligence. This position paper charts the trajectory of this emerging field, arguing that it is evolving from automating…

计算与语言 · 计算机科学 2026-01-30 Yiran Rex Ma

Chinese pre-trained language models usually process text as a sequence of characters, while ignoring more coarse granularity, e.g., words. In this work, we propose a novel pre-training paradigm for Chinese -- Lattice-BERT, which explicitly…

计算与语言 · 计算机科学 2021-05-31 Yuxuan Lai , Yijia Liu , Yansong Feng , Songfang Huang , Dongyan Zhao

The rapid growth of visual tokens in multimodal large language models (MLLMs) leads to excessive memory consumption and inference latency, especially when handling high-resolution images and videos. Token pruning is a technique used to…

计算机视觉与模式识别 · 计算机科学 2025-12-02 Zhongyu Yang , Dannong Xu , Wei Pang , Yingfang Yuan

Most existing online writer-identification systems require that the text content is supplied in advance and rely on separately designed features and classifiers. The identifications are based on lines of text, entire paragraphs, or entire…

计算机视觉与模式识别 · 计算机科学 2015-05-20 Weixin Yang , Lianwen Jin , Manfei Liu

The archaeological dating of bronze dings has played a critical role in the study of ancient Chinese history. Current archaeology depends on trained experts to carry out bronze dating, which is time-consuming and labor-intensive. For such…

计算机视觉与模式识别 · 计算机科学 2023-06-05 Rixin Zhou , Jiafu Wei , Qian Zhang , Ruihua Qi , Xi Yang , Chuntao Li

Bronze inscriptions (BI), engraved on ritual vessels, constitute a crucial stage of early Chinese writing and provide indispensable evidence for archaeological and historical studies. However, automatic BI recognition remains difficult due…

计算机视觉与模式识别 · 计算机科学 2025-10-03 Rixin Zhou , Peiqiang Qiu , Qian Zhang , Chuntao Li , Xi Yang

The information provided by historical documents has always been indispensable in the transmission of human civilization, but it has also made these books susceptible to damage due to various factors. Thanks to recent technology, the…

计算机视觉与模式识别 · 计算机科学 2021-04-06 Chia-Wei Tang , Chao-Lin Liu , Po-Sen Chiu

Recently, word enhancement has become very popular for Chinese Named Entity Recognition (NER), reducing segmentation errors and increasing the semantic and boundary information of Chinese words. However, these methods tend to ignore the…

计算与语言 · 计算机科学 2021-07-13 Shuang Wu , Xiaoning Song , Zhenhua Feng

Pretrained language models (PLMs) have shown marvelous improvements across various NLP tasks. Most Chinese PLMs simply treat an input text as a sequence of characters, and completely ignore word information. Although Whole Word Masking can…

计算与语言 · 计算机科学 2023-03-23 Xinnian Liang , Zefan Zhou , Hui Huang , Shuangzhi Wu , Tong Xiao , Muyun Yang , Zhoujun Li , Chao Bian

Chinese is one of the most widely used languages in the world, yet online handwritten Chinese character recognition (OLHCCR) remains challenging. To recognize Chinese characters, one popular choice is to adopt the 2D convolutional neural…

计算机视觉与模式识别 · 计算机科学 2020-04-21 Ji Gan , Weiqiang Wang , Ke Lu

Given the advantage and recent success of English character-level and subword-unit models in several NLP tasks, we consider the equivalent modeling problem for Chinese. Chinese script is logographic and many Chinese logograms are composed…

计算与语言 · 计算机科学 2018-09-11 Falcon Z. Dai , Zheng Cai

We present TMMLU+, a new benchmark designed for Traditional Chinese language understanding. TMMLU+ is a multi-choice question-answering dataset with 66 subjects from elementary to professional level. It is six times larger and boasts a more…

计算与语言 · 计算机科学 2024-07-12 Zhi-Rui Tam , Ya-Ting Pai , Yen-Wei Lee , Jun-Da Chen , Wei-Min Chu , Sega Cheng , Hong-Han Shuai

Identifying the different varieties of the same language is more challenging than unrelated languages identification. In this paper, we propose an approach to discriminate language varieties or dialects of Mandarin Chinese for the Mainland…

计算与语言 · 计算机科学 2017-01-10 Fan Xu , Mingwen Wang , Maoxi Li

Ancient Chinese word segmentation (WSG) and part-of-speech tagging (POS) are important to study ancient Chinese, but the amount of ancient Chinese WSG and POS tagging data is still rare. In this paper, we propose a novel augmentation method…

计算与语言 · 计算机科学 2023-03-07 Shuo Feng , Piji Li

Disambiguating scholars with identical names is essential for accurate authorship assignment and robust large-scale scientometric research. Existing methods are often designed for Latin-script metadata and perform poorly on Chinese names.…

数字图书馆 · 计算机科学 2026-04-07 Mingrong She , Liuhuaying Yang , Ana Maria Jaramillo , Lisette Espín-Noboa

Oracle bone script, one of the earliest known forms of ancient Chinese writing, presents invaluable research materials for scholars studying the humanities and geography of the Shang Dynasty, dating back 3,000 years. The immense historical…

计算机视觉与模式识别 · 计算机科学 2024-09-04 Pengjie Wang , Kaile Zhang , Xinyu Wang , Shengwei Han , Yongge Liu , Jinpeng Wan , Haisu Guan , Zhebin Kuang , Lianwen Jin , Xiang Bai , Yuliang Liu

Multi-stroke characters in scripts such as Chinese and Japanese can be highly complex, posing significant challenges for both native speakers and, especially, non-native learners. If these characters can be simplified without degrading…

计算机视觉与模式识别 · 计算机科学 2025-07-01 Ryo Ishiyama , Shinnosuke Matsuo , Seiichi Uchida

Recent pretraining models in Chinese neglect two important aspects specific to the Chinese language: glyph and pinyin, which carry significant syntax and semantic information for language understanding. In this work, we propose ChineseBERT,…

计算与语言 · 计算机科学 2021-07-01 Zijun Sun , Xiaoya Li , Xiaofei Sun , Yuxian Meng , Xiang Ao , Qing He , Fei Wu , Jiwei Li