中文
相关论文

相关论文: Non-Standard Words as Features for Text Categoriza…

200 篇论文

Document categorization is a technique where the category of a document is determined. In this paper three well-known supervised learning techniques which are Support Vector Machine(SVM), Na\"ive Bayes(NB) and Stochastic Gradient…

计算与语言 · 计算机科学 2017-01-31 Md. Saiful Islam , Fazla Elahi Md Jubayer , Syed Ikhtiar Ahmed

The factor complexity function $C_w(n)$ of a finite or infinite word $w$ counts the number of distinct factors of $w$ of length $n$ for each $n \ge 0$. A finite word $w$ of length $|w|$ is said to be trapezoidal if the graph of its factor…

组合数学 · 数学 2015-02-25 Amy Glen , Florence Levé

Croatian is poorly resourced and highly inflected language from Slavic language family. Nowadays, research is focusing mostly on English. We created a new word analogy corpus based on the original English Word2vec word analogy corpus and…

计算与语言 · 计算机科学 2017-11-09 Lukas Svoboda , Slobodan Beliga

In text processing, deep neural networks mostly use word embeddings as an input. Embeddings have to ensure that relations between words are reflected through distances in a high-dimensional numeric space. To compare the quality of different…

计算与语言 · 计算机科学 2022-06-01 Matej Ulčar , Kristiina Vaik , Jessica Lindström , Milda Dailidėnaitė , Marko Robnik-Šikonja

Slot filling is an important problem in Spoken Language Understanding (SLU) and Natural Language Processing (NLP), which involves identifying a user's intent and assigning a semantic concept to each word in a sentence. This paper presents a…

计算与语言 · 计算机科学 2018-06-20 Ruixi Lin

Several methods have been explored for automating parts of Systematic Mapping (SM) and Systematic Review (SR) methodologies. Challenges typically evolve around the gaps in semantic understanding of text, as well as lack of domain and…

计算与语言 · 计算机科学 2021-02-10 Xiajing Li , Marios Daoutis

Word similarity has many applications to social science and cultural analytics tasks like measuring meaning change over time and making sense of contested terms. Yet traditional similarity methods based on cosine similarity between word…

计算与语言 · 计算机科学 2025-02-11 Kaitlyn Zhou , Haishan Gao , Sarah Chen , Dan Edelstein , Dan Jurafsky , Chen Shani

Word sense disambiguation (WSD) is a long-standing problem in natural language processing. One significant challenge in supervised all-words WSD is to classify among senses for a majority of words that lie in the long-tail distribution. For…

计算与语言 · 计算机科学 2021-04-28 Howard Chen , Mengzhou Xia , Danqi Chen

This paper presents a collection of highly comparable web corpora of Slovenian, Croatian, Bosnian, Montenegrin, Serbian, Macedonian, and Bulgarian, covering thereby the whole spectrum of official languages in the South Slavic language…

计算与语言 · 计算机科学 2024-05-28 Nikola Ljubešić , Taja Kuzman

Text Classification is the process of categorizing text into the relevant categories and its algorithms are at the core of many Natural Language Processing (NLP). Term Frequency-Inverse Document Frequency (TF-IDF) and NLP are the most…

计算与语言 · 计算机科学 2023-08-09 Mamata Das , Selvakumar K. , P. J. A. Alphonse

Automating primary stress identification has been an active research field due to the role of stress in encoding meaning and aiding speech comprehension. Previous studies relied mainly on traditional acoustic features and English datasets.…

音频与语音处理 · 电气工程与系统科学 2025-06-02 Nikola Ljubešić , Ivan Porupski , Peter Rupnik

Clustering a lexicon of words is a well-studied problem in natural language processing (NLP). Word clusters are used to deal with sparse data in statistical language processing, as well as features for solving various NLP tasks (text…

计算与语言 · 计算机科学 2018-08-17 Effi Levi , Saggy Herman , Ari Rappoport

This paper presents a new training dataset for automatic genre identification GINCO, which is based on 1,125 crawled Slovenian web documents that consist of 650 thousand words. Each document was manually annotated for genre with a new…

计算与语言 · 计算机科学 2022-01-12 Taja Kuzman , Peter Rupnik , Nikola Ljubešić

Statistical methods have been widely employed in recent years to grasp many language properties. The application of such techniques have allowed an improvement of several linguistic applications, which encompasses machine translation,…

计算与语言 · 计算机科学 2016-02-22 Henrique F. de Arruda , Luciano da F. Costa , Diego R. Amancio

This research introduces a novel psychometric method for analyzing textual data using large language models. By leveraging contextual embeddings to create contextual scores, we transform textual data into response data suitable for…

计算与语言 · 计算机科学 2025-09-12 Jinsong Chen

Since Bahdanau et al. [1] first introduced attention for neural machine translation, most sequence-to-sequence models made use of attention mechanisms [2, 3, 4]. While they produce soft-alignment matrices that could be interpreted as…

计算与语言 · 计算机科学 2019-09-12 Marcely Zanon Boito , Aline Villavicencio , Laurent Besacier

Text document classification is an important task for diverse natural language processing based applications. Traditional machine learning approaches mainly focused on reducing dimensionality of textual data to perform classification. This…

Enhancing word usage is a desired feature for writing assistance. To further advance research in this area, this paper introduces "Smart Word Suggestions" (SWS) task and benchmark. Unlike other works, SWS emphasizes end-to-end evaluation…

计算与语言 · 计算机科学 2023-05-18 Chenshuo Wang , Shaoguang Mao , Tao Ge , Wenshan Wu , Xun Wang , Yan Xia , Jonathan Tien , Dongyan Zhao

Classifying text is a method for categorizing documents into pre-established groups. Text documents must be prepared and represented in a way that is appropriate for the algorithms used for data mining prior to classification. As a result,…

计算与语言 · 计算机科学 2024-02-26 Esra'a Alhenawi , Ruba Abu Khurma , Pedro A. Castillo , Maribel G. Arenas

In this paper we analyse network motifs in the co-occurrence directed networks constructed from five different texts (four books and one portal) in the Croatian language. After preparing the data and network construction, we perform the…

计算与语言 · 计算机科学 2014-11-19 Hana Rizvić , Sanda Martinčić-Ipšić , Ana Meštrović