English
Related papers

Related papers: Quantifying patterns of punctuation in modern Chin…

200 papers

A statistical model for describing the scaling of the distribution of inter-event times is described. By considering the diverse region seismicity (natural and induced) at different scale levels the self-similarity of the distribution has…

Other Condensed Matter · Physics 2009-11-11 V. German

Text style transfer task requires the model to transfer a sentence of one style to another style while retaining its original content meaning, which is a challenging problem that has long suffered from the shortage of parallel data. In this…

Computation and Language · Computer Science 2019-09-26 Mingyue Shang , Piji Li , Zhenxin Fu , Lidong Bing , Dongyan Zhao , Shuming Shi , Rui Yan

We investigate a lattice LSTM network for Chinese word segmentation (CWS) to utilize words or subwords. It integrates the character sequence features with all subsequences information matched from a lexicon. The matched subsequences serve…

Computation and Language · Computer Science 2018-10-31 Jie Yang , Yue Zhang , Shuailong Liang

Zipf's law can be used to describe the rank-size distribution of cities in a region. It was seldom employed to research urban internal structure. In this paper, we demonstrate that the space-filling process within a city follows Zipf's law…

Physics and Society · Physics 2018-12-19 Yanguang Chen , Jiejing Wang

Distributed word representations are very useful for capturing semantic information and have been successfully applied in a variety of NLP tasks, especially on English. In this work, we innovatively develop two component-enhanced Chinese…

Computation and Language · Computer Science 2015-08-28 Yanran Li , Wenjie Li , Fei Sun , Sujian Li

Chinese word segmentation (CWS) is often regarded as a character-based sequence labeling task in most current works which have achieved great success with the help of powerful neural networks. However, these works neglect an important clue:…

Computation and Language · Computer Science 2019-05-31 Jingkang Wang , Jianing Zhou , Jie Zhou , Gongshen Liu

Long-range correlations are found in symbolic sequences from human language, music and DNA. Determining the span of correlations in dolphin whistle sequences is crucial for shedding light on their communicative complexity. Dolphin whistles…

Neurons and Cognition · Quantitative Biology 2014-12-03 Ramon Ferrer-i-Cancho , Brenda McCowan

The distribution of the lifetime of Chinese dynasties (as well as that of the British Isles and Japan) in a linear Zipf plot is found to consist of two straight lines intersecting at a transition point. This two-section piecewise-linear…

In this paper we explore where information is collected and how it is propagated throughout layers in large language models (LLMs). We begin by examining the surprising computational importance of punctuation tokens which previous work has…

Computation and Language · Computer Science 2025-08-21 Sonakshi Chauhan , Maheep Chaudhary , Koby Choy , Samuel Nellessen , Nandi Schoots

The Weibull function is widely used to describe skew distributions observed in nature. However, the origin of this ubiquity is not always obvious to explain. In the present paper, we consider the well-known Galton-Watson branching process…

Data Analysis, Statistics and Probability · Physics 2015-05-27 Junghyo Jo , Jean-Yves Fortin , M. Y. Choi

The task of text segmentation may be undertaken at many levels in text analysis---paragraphs, sentences, words, or even letters. Here, we focus on a relatively fine scale of segmentation, hypothesizing it to be in accord with a stochastic…

Pretrained language models (PLMs) have shown marvelous improvements across various NLP tasks. Most Chinese PLMs simply treat an input text as a sequence of characters, and completely ignore word information. Although Whole Word Masking can…

Computation and Language · Computer Science 2023-03-23 Xinnian Liang , Zefan Zhou , Hui Huang , Shuangzhi Wu , Tong Xiao , Muyun Yang , Zhoujun Li , Chao Bian

We propose a new approach to the Chinese word segmentation problem that considers the sentence as an undirected graph, whose nodes are the characters. One can use various techniques to compute the edge weights that measure the connection…

Computation and Language · Computer Science 2018-04-06 Yuanhao Liu , Sheng Yu

Moral sentiments expressed in natural language significantly influence both online and offline environments, shaping behavioral styles and interaction patterns, including social media selfpresentation, cyberbullying, adherence to social…

Computation and Language · Computer Science 2024-11-15 Renjie Cao , Miaoyan Hu , Jiahan Wei , Baha Ihnaini

Outline generation aims to reveal the internal structure of a document by identifying underlying chapter relationships and generating corresponding chapter summaries. Although existing deep learning methods and large models perform well on…

Artificial Intelligence · Computer Science 2024-12-03 Yan Yan , Yuanchi Ma

Due to the availability of references of research papers and the rich information contained in papers, various citation analysis approaches have been proposed to identify similar documents for scholar recommendation. Despite of the success…

Information Retrieval · Computer Science 2017-03-21 Han Tian , Hankz Hankui Zhuo

Measures of textual similarity and divergence are increasingly used to study cultural change. But which measures align, in practice, with social evidence about change? We apply three different representations of text (topic models, document…

Computation and Language · Computer Science 2024-11-25 Sarah Griebel , Becca Cohen , Lucian Li , Jaihyun Park , Jiayu Liu , Jana Perkins , Ted Underwood

Ancient Chinese word segmentation (WSG) and part-of-speech tagging (POS) are important to study ancient Chinese, but the amount of ancient Chinese WSG and POS tagging data is still rare. In this paper, we propose a novel augmentation method…

Computation and Language · Computer Science 2023-03-07 Shuo Feng , Piji Li

Chinese pre-trained language models usually process text as a sequence of characters, while ignoring more coarse granularity, e.g., words. In this work, we propose a novel pre-training paradigm for Chinese -- Lattice-BERT, which explicitly…

Computation and Language · Computer Science 2021-05-31 Yuxuan Lai , Yijia Liu , Yansong Feng , Songfang Huang , Dongyan Zhao

Starting from a master equation, we derive the evolution equation for the size distribution of elements in an evolving system, where each element can grow, divide into two, and produce new elements. We then probe general solutions of the…

Statistical Mechanics · Physics 2011-04-05 Segun Goh , H. W. Kwon , M. Y. Choi , Jean-Yves Fortin