English
Related papers

Related papers: Quantifying patterns of punctuation in modern Chin…

200 papers

Most of the Chinese pre-trained models adopt characters as basic units for downstream tasks. However, these models ignore the information carried by words and thus lead to the loss of some important semantics. In this paper, we propose a…

Computation and Language · Computer Science 2022-07-14 Wenbiao Li , Rui Sun , Yunfang Wu

Chinese Spelling Correction (CSC) aims to detect and correct spelling errors in Chinese sentences caused by phonetic or visual similarities. While current CSC models integrate pinyin or glyph features and have shown significant…

Computation and Language · Computer Science 2024-09-10 Lei Sheng , Shuai-Shuai Xu

Language, which allows complex ideas to be communicated through symbolic sequences, is a characteristic feature of our species and manifested in a multitude of forms. Using large written corpora for many different languages and scripts, we…

Computation and Language · Computer Science 2018-01-17 Md Izhar Ashraf , Sitabhra Sinha

We study the entropy of Chinese and English texts, based on characters in case of Chinese texts and based on words for both languages. Significant differences are found between the languages and between different personal styles of debating…

Computation and Language · Computer Science 2017-01-17 R. R. Xie , W. B. Deng , D. J. Wang , L. P. Csernai

Universality or near-universality of citation distributions was found empirically a decade ago but its theoretical justification has been lacking so far. Here, we systematically study citation distributions for different disciplines in…

Physics and Society · Physics 2020-02-25 Michael Golosovsky

Zipf's law seems to be ubiquitous in human languages and appears to be a universal property of complex communicating systems. Following the early proposal made by Zipf concerning the presence of a tension between the efforts of speaker and…

Adaptation and Self-Organizing Systems · Physics 2015-05-19 Bernat Corominas-Murtra , Jordi Fortuny , Ricard V. Solé

In this paper we address a method to align English-Chinese bilingual news reports from China News Service, combining both lexical and satistical approaches. Because of the sentential structure differences between English and Chinese,…

cmp-lg · Computer Science 2008-02-03 Donghua Xu , Chew Lim Tan

Inspired by early research on exploring naturally annotated data for Chinese word segmentation (CWS), and also by recent research on integration of speech and text processing, this work for the first time proposes to mine word boundaries…

Computation and Language · Computer Science 2023-10-31 Lei Zhang , Zhenghua Li , Shilin Zhou , Chen Gong , Zhefeng Wang , Baoxing Huai , Min Zhang

In the case of the scientometric evaluation of multi- or interdisciplinary units one risks to compare apples with oranges: each paper has to assessed in comparison to an appropriate reference set. We suggest that the set of citing papers…

Digital Libraries · Computer Science 2010-12-03 Ping Zhou , Loet Leydesdorff

Burrows' Delta was introduced in 2002 and has proven to be an effective tool for author attribution. Despite the fact that these are different languages, they mostly belong to the same grammatical type and use the same graphic principle to…

Computation and Language · Computer Science 2024-07-12 Boris Orekhov

Sentence Simplification is a valuable technique that can benefit language learners and children a lot. However, current research focuses more on English sentence simplification. The development of Chinese sentence simplification is…

Computation and Language · Computer Science 2023-06-08 Shiping Yang , Renliang Sun , Xiaojun Wan

Authorship attribution models fine-tuned with the same pretrained encoder, data, and loss can differ four-fold in performance depending only on their scoring mechanism. We use mechanistic interpretability tools to explain this gap.…

Computation and Language · Computer Science 2026-05-27 Francis Kulumba , Guillaume Vimont , Laurent Romary , Florian Cafiero

Recent observations in the theory of verse and empirical metrics have suggested that constructing a verse line involves a pattern-matching search through a source text, and that the number of found elements (complete words totaling a…

cmp-lg · Computer Science 2007-05-23 Hideaki Aoyama , John Constable

We are developing a knowledge base over Chinese judicial decision documents to facilitate landscape analyses of Chinese Criminal Cases. We view judicial decision documents as a mixed-granularity semi-structured text where different levels…

Computers and Society · Computer Science 2019-10-17 Xiaohan Wu , Benjamin L. Liebman , Rachel E. Stern , Margaret E. Roberts , Amarnath Gupta

Multi-Criteria Chinese Word Segmentation (MCCWS) aims at finding word boundaries in a Chinese sentence composed of continuous characters while multiple segmentation criteria exist. The unified framework has been widely used in MCCWS and…

Computation and Language · Computer Science 2020-04-14 Zhen Ke , Liang Shi , Erli Meng , Bin Wang , Xipeng Qiu , Xuanjing Huang

In this article, the relationship between two well-accepted empirical propositions regarding the distribution of population in cities, namely, Gibrat's law and Zipf's law, are rigorously examined using the Chinese census data. Our findings…

Physics and Society · Physics 2010-01-07 Kausik Gangopadhyay , B. Basu

In constituency parsing, span-based decoding is an important direction. However, for Chinese sentences, because of their linguistic characteristics, it is necessary to utilize other models to perform word segmentation first, which…

Computation and Language · Computer Science 2022-12-01 Zhicheng Wang , Tianyu Shi , Cong Liu

We use the formulation of equilibrium statistical mechanics in order to study some important characteristics of language. Using a simple expression for the Hamiltonian of a language system, which is directly implied by the Zipf law, we are…

Physics and Society · Physics 2009-11-11 Kosmas Kosmidis , Alkiviadis Kalampokis , Panos Argyrakis

A realistic Chinese word segmentation tool must adapt to textual variations with minimal training input and yet robust enough to yield reliable segmentation result for all variants. Various lexicon-driven approaches to Chinese segmentation,…

Computation and Language · Computer Science 2019-05-22 Chu-Ren Huang , Ting-Shuo Yo , Petr Simon , Shu-Kai Hsieh

We consider three major text sources about the Tang Dynasty of China in our experiments that aim to segment text written in classical Chinese. These corpora include a collection of Tang Tomb Biographies, the New Tang Book, and the Old Tang…

Computation and Language · Computer Science 2020-07-23 Chao-Lin Liu , Chang-Ting Chu , Wei-Ting Chang , Ti-Yong Zheng
‹ Prev 1 4 5 6 7 8 10 Next ›