中文
相关论文

相关论文: Non-Standard Words as Features for Text Categoriza…

200 篇论文

This paper presents text normalization which is an integral part of any text-to-speech synthesis system. Text normalization is a set of methods with a task to write non-standard words, like numbers, dates, times, abbreviations, acronyms and…

计算与语言 · 计算机科学 2015-03-31 Slobodan Beliga , Miran Pobar , Sanda Martinčić-Ipšić

Text Normalization is an integral part of any text-to-speech synthesis system. In a natural language text, there are elements such as numbers, dates, abbreviations, etc. that belong to other semiotic classes. They are called non-standard…

计算与语言 · 计算机科学 2024-01-02 Pijus Kasparaitis

Converting written texts into their spoken forms is an essential problem in any text-to-speech (TTS) systems. However, building an effective text normalization solution for a real-world TTS system face two main challenges: (1) the semantic…

计算与语言 · 计算机科学 2022-09-08 Huu-Tien Dang , Thi-Hai-Yen Vuong , Xuan-Hieu Phan

In this paper, we use statistical texture features for handwritten and printed text classification. We primarily aim for word level classification in south Indian scripts. Words are first extracted from the scanned document. For each…

计算机视觉与模式识别 · 计算机科学 2013-04-11 Mallikarjun Hangarge , K. C. Santosh , Srikanth Doddamani , Rajmohan Pardeshi

Text Categorization is traditionally done by using the term frequency and inverse document frequency.This type of method is not very good because, some words which are not so important may appear in the document .The term frequency of…

信息检索 · 计算机科学 2016-11-25 Srikanth Bethu , G Charless Babu , J Vinoda , E Priyadarshini , M Raghavendra rao

In this work we analyze statistical properties of 91 relatively small texts in 7 different languages (Spanish, English, French, German, Turkish, Russian, Icelandic) as well as texts with randomly inserted spaces. Despite the size (around…

物理与社会 · 物理学 2020-07-15 Diego Espitia , Hernán Larralde

Spell-checkers are valuable tools that enhance communication by identifying misspelled words in written texts. Recent improvements in deep learning, and in particular in large language models, have opened new opportunities to improve…

计算与语言 · 计算机科学 2024-11-01 Matej Klemen , Martin Božič , Špela Arhar Holdt , Marko Robnik-Šikonja

Preliminary report on network based keyword extraction for Croatian is an unsupervised method for keyword extraction from the complex network. We build our approach with a new network measure the node selectivity, motivated by the research…

计算与语言 · 计算机科学 2018-02-15 Slobodan Beliga , Ana Meštrović , Sanda Martinčić-Ipšić

Volume of text based documents have been increasing day by day. Medical documents are located within this growing text documents. In this study, the techniques used for text classification applied on medical documents and evaluated…

信息检索 · 计算机科学 2018-11-05 Durmus Ozkan Sahin , Erdal Kilic

Language identification is an important Natural Language Processing task. It has been thoroughly researched in the literature. However, some issues are still open. This work addresses the identification of the related low-resource languages…

计算与语言 · 计算机科学 2022-03-10 Olha Dovbnia , Anna Wróblewska

This paper describes the National Research Council (NRC) Word Sense Disambiguation (WSD) system, as applied to the English Lexical Sample (ELS) task in Senseval-3. The NRC system approaches WSD as a classical supervised machine learning…

计算与语言 · 计算机科学 2007-05-23 Peter D. Turney

In this paper, we focus on the detection of semantic changes in Slovene, a less resourced Slavic language with two million speakers. Detecting and tracking semantic changes provides insight into the evolution of language caused by changes…

计算与语言 · 计算机科学 2025-05-29 Marko Pranjić , Kaja Dobrovoljc , Senja Pollak , Matej Martinc

The Ord's graph is a simple graphical method for displaying frequency distributions of data or theoretical distributions in the two-dimensional plane. Its coordinates are proportions of the first three moments, either empirical or…

应用统计 · 统计学 2015-04-15 Michaela Koscová , Ján Macutek , Emmerich Kelih

Many of the existing TTS systems cannot accurately synthesize text containing a variety of numerical formats, resulting in reduced intelligibility of the synthesized speech. This research aims to develop a numerical format classifier that…

音频与语音处理 · 电气工程与系统科学 2024-12-03 Yaser Darwesh , Lit Wei Wern , Mumtaz Begum Mustafa

Text classification plays a vital role today especially with the intensive use of social networking media. Recently, different architectures of convolutional neural networks have been used for text classification in which one-hot vector,…

计算与语言 · 计算机科学 2019-03-12 Amr Adel Helmy , Yasser M. K. Omar , Rania Hodhod

Word feature vectors have been proven to improve many NLP tasks. With recent advances in unsupervised learning of these feature vectors, it became possible to train it with much more data, which also resulted in better quality of learned…

计算与语言 · 计算机科学 2022-11-29 Marius Sajgalik , Michal Barla , Maria Bielikova

Measuring how semantics of words change over time improves our understanding of how cultures and perspectives change. Diachronic word embeddings help us quantify this shift, although previous studies leveraged substantial temporally…

计算与语言 · 计算机科学 2025-06-17 David Dukić , Ana Barić , Marko Čuljak , Josip Jukić , Martin Tutek

In stylometric investigations, frequencies of the most frequent words (MFWs) and character n-grams outperform other style-markers, even if their performance varies significantly across languages. In inflected languages, word endings play a…

计算与语言 · 计算机科学 2022-11-03 Maciej Eder , Rafał. L. Górski

This paper presents a dataset and supervised learning experiments for term extraction from Slovene academic texts. Term candidates in the dataset were extracted via morphosyntactic patterns and annotated for their termness by four…

计算与语言 · 计算机科学 2019-06-06 Nikola Ljubešić , Darja Fišer , Tomaž Erjavec

Using NLP to analyze authentic learner language helps to build automated assessment and feedback tools. It also offers new and extensive insights into the development of second language production. However, there is a lack of research…

计算与语言 · 计算机科学 2026-03-10 Kais Allkivi
‹ 上一页 1 2 3 10 下一页 ›