中文
相关论文

相关论文: Non-Standard Words as Features for Text Categoriza…

200 篇论文

This paper (cmp-lg/yymmnnn) has been accepted for publication in the student session of EACL-95. It outlines ongoing work using statistical and unsupervised neural network methods for clustering words in untagged corpora. Such approaches…

cmp-lg · 计算机科学 2008-02-03 Christopher C. Huckle

We study a deliberately simple, fully non-linguistic model of text: a sequence of independent draws from a finite alphabet of letters plus a single space symbol. A word is defined as a maximal block of non-space symbols. Within this…

计算与语言 · 计算机科学 2025-11-25 Vladimir Berman

Text classification has become indispensable due to the rapid increase of text in digital form. Over the past three decades, efforts have been made to approach this task using various learning algorithms and statistical models based on…

机器学习 · 统计学 2018-06-11 Erica K. Shimomoto , Lincon S. Souza , Bernardo B. Gatto , Kazuhiro Fukui

Character recognition techniques for printed documents are widely used for English language. However, the systems that are implemented to recognize Asian languages struggle to increase the accuracy of recognition. Among other Asian…

计算机视觉与模式识别 · 计算机科学 2014-12-25 G. I. Gunarathna , M. A. P. Chamikara , R. G. Ragel

This work aims to automatically evaluate whether the language development of children is age-appropriate. Validated speech and language tests are used for this purpose to test the auditory memory. In this work, the task is to determine…

音频与语音处理 · 电气工程与系统科学 2022-06-20 Ilja Baumann , Dominik Wagner , Sebastian Bayerl , Tobias Bocklet

Speech Acts (SAs) are one of the important areas of pragmatics, which give us a better understanding of the state of mind of the people and convey an intended language function. Knowledge of the SA of a text can be helpful in analyzing that…

计算与语言 · 计算机科学 2020-07-14 Zoleikha Jahanbakhsh-Nagadeh , Mohammad-Reza Feizi-Derakhshi , Arash Sharifi

Since the seminal work of Mikolov et al., word embeddings have become the preferred word representations for many natural language processing tasks. Document similarity measures extracted from word embeddings, such as the soft cosine…

信息检索 · 计算机科学 2020-04-02 Vít Novotný , Eniafe Festus Ayetiran , Michal Štefánik , Petr Sojka

Natural Language Inference (NLI) and Semantic Textual Similarity (STS) are widely used benchmark tasks for compositional evaluation of pre-trained language models. Despite growing interest in linguistic universals, most NLI/STS studies have…

计算与语言 · 计算机科学 2022-08-10 Hitomi Yanaka , Koji Mineshima

The writing style of a person can be affirmed as a unique identity indicator; the words used, and the structuring of the sentences are clear measures which can identify the author of a specific work. Stylometry and its subset - Authorship…

计算与语言 · 计算机科学 2018-12-27 Abhay Sharma , Ananya Nandan , Reetika Ralhan

Local/Native South African languages are classified as low-resource languages. As such, it is essential to build the resources for these languages so that they can benefit from advances in the field of natural language processing. In this…

计算与语言 · 计算机科学 2023-06-14 Andani Madodonga , Vukosi Marivate , Matthew Adendorff

Short text clustering is a known use case in the text analytics community. When the structure and content falls in the natural language domain e.g. Twitter posts or instant messages, then natural language techniques can be used, provided…

机器学习 · 计算机科学 2025-09-01 Thanasis Schoinas , Benjamin Guinard , Diba Esbati , Richard Chalk

A well-known but rarely used approach to text categorization uses conditional entropy estimates computed using data compression tools. Text affinity scores derived from compressed sizes can be used for classification and ranking tasks, but…

机器学习 · 计算机科学 2021-12-08 Nitya Kasturi , Igor L. Markov

Much work has been done on feature selection. Existing methods are based on document frequency, such as Chi-Square Statistic, Information Gain etc. However, these methods have two shortcomings: one is that they are not reliable for…

机器学习 · 计算机科学 2013-05-06 Deqing Wang , Hui Zhang , Rui Liu , Weifeng Lv

Given a random text over a finite alphabet, we study the frequencies at which fixed-length words occur as subsequences. As the data size grows, the joint distribution of word counts exhibits a rich asymptotic structure. We investigate all…

概率论 · 数学 2026-05-06 Chaim Even-Zohar , Tsviqa Lakrec , Ran J. Tessler

With such increasing popularity and availability of digital text data, authorships of digital texts can not be taken for granted due to the ease of copying and parsing. This paper presents a new text style analysis called natural frequency…

计算与语言 · 计算机科学 2012-08-16 Zhili Chen , Liusheng Huang , Wei Yang , Peng Meng , Haibo Miao

Objective: To detect and classify features of stigmatizing and biased language in intensive care electronic health records (EHRs) using natural language processing techniques. Materials and Methods: We first created a lexicon and regular…

计算与语言 · 计算机科学 2025-07-15 Drew Walker , Annie Thorne , Sudeshna Das , Jennifer Love , Hannah LF Cooper , Melvin Livingston , Abeed Sarker

Lemmatization is a Natural Language Processing (NLP) technique used to normalize text by changing morphological derivations of words to their root forms. It is used as a core pre-processing step in many NLP tasks including text indexing,…

计算与语言 · 计算机科学 2023-08-04 Shafie Abdi Mohamed , Muhidin Abdullahi Mohamed

The deviation of the observed frequency of a word $w$ from its expected frequency in a given sequence $x$ is used to determine whether or not the word is avoided. This concept is particularly useful in DNA linguistic analysis. The value of…

For sensible progress in natural language processing, it is important that we are aware of the limitations of the evaluation metrics we use. In this work, we evaluate how robust metrics are to non-standardized dialects, i.e. spelling…

计算与语言 · 计算机科学 2023-11-29 Noëmi Aepli , Chantal Amrhein , Florian Schottmann , Rico Sennrich

Central, standard, and Christoffel words are three strongly interrelated classes of binary finite words which represent a finite counterpart of characteristic Sturmian words. A natural arithmetization of the theory is obtained by…

离散数学 · 计算机科学 2014-10-16 Aldo de Luca , Alessandro De Luca