中文
相关论文

相关论文: Non-Standard Words as Features for Text Categoriza…

200 篇论文

Time series classification is an application of particular interest with the increase of data to monitor. Classical techniques for time series classification rely on point-to-point distances. Recently, Bag-of-Words approaches have been used…

机器学习 · 计算机科学 2016-01-14 Adeline Bailly , Simon Malinowski , Romain Tavenard , Thomas Guyet , Laetitia Chapel

We characterize the clustering of a word under the Burrows-Wheeler transform in terms of the resolution of a bounded number of bispecial factors belonging to the language generated by all its powers. We use this criterion to compute, in…

动力系统 · 数学 2023-05-31 Sébastien Ferenczi , Luca Q. Zamboni

Abbreviations present a significant challenge for NLP systems because they cause tokenization and out-of-vocabulary errors. They can also make the text less readable, especially in reference printed books, where they are extensively used.…

计算与语言 · 计算机科学 2022-11-07 Angel Daza , Antske Fokkens , Tomaž Erjavec

Survey data can contain a high number of features while having a comparatively low quantity of examples. Machine learning models that attempt to predict outcomes from survey data under these conditions can overfit and result in poor…

计算与语言 · 计算机科学 2023-08-22 Benjamin C. Warner , Ziqi Xu , Simon Haroutounian , Thomas Kannampallil , Chenyang Lu

Semantic Shift Detection (SSD) is the task of identifying, interpreting, and assessing the possible change over time in the meanings of a target word. Traditionally, SSD has been addressed by linguists and social scientists through manual…

计算与语言 · 计算机科学 2024-06-12 Stefano Montanelli , Francesco Periti

Every speech signal carries implicit information about the emotions, which can be extracted by speech processing methods. In this paper, we propose an algorithm for extracting features that are independent from the spoken language and the…

音频与语音处理 · 电气工程与系统科学 2018-11-26 Fatemeh Noroozi , Marina Marjanovic , Angelina Njegus , Sergio Escalera , Gholamreza Anbarjafari

This article describes the results of a systematic in-depth study of the criteria used for word sense disambiguation. Our study is based on 60 target words: 20 nouns, 20 adjectives and 20 verbs. Our results are not always in line with some…

计算与语言 · 计算机科学 2007-05-23 Laurent Audibert

Recently, the supervised learning paradigm's surprisingly remarkable performance has garnered considerable attention from Sanskrit Computational Linguists. As a result, the Sanskrit community has put laudable efforts to build task-specific…

计算与语言 · 计算机科学 2021-04-02 Jivnesh Sandhan , Om Adideva , Digumarthi Komal , Laxmidhar Behera , Pawan Goyal

Separable Non-negative Matrix Factorization (SNMF) is an important method for topic modeling, where "separable" assumes every topic contains at least one anchor word, defined as a word that has non-zero probability only on that topic. SNMF…

信息检索 · 计算机科学 2019-05-16 Kun He , Wu Wang , Xiaosen Wang , John E. Hopcroft

The referential properties of noun phrases in the Japanese language, which has no articles, are useful for article generation in Japanese-English machine translation and for anaphora resolution in Japanese noun phrases. They are generally…

计算与语言 · 计算机科学 2007-05-23 Masaki Murata , Kiyotaka Uchimoto , Qing Ma , Hitoshi Isahara

Feature extraction is an important process of machine learning and deep learning, as the process make algorithms function more efficiently, and also accurate. In natural language processing used in deception detection such as fake news…

计算与语言 · 计算机科学 2020-11-04 HyeonJun Kim

The goal of this work is to design a machine translation (MT) system for a low-resource family of dialects, collectively known as Swiss German, which are widely spoken in Switzerland but seldom written. We collected a significant number of…

计算与语言 · 计算机科学 2018-02-07 Pierre-Edouard Honnet , Andrei Popescu-Belis , Claudiu Musat , Michael Baeriswyl

This study investigates global properties of literary and non-literary texts. Within the literary texts, a distinction is made between canonical and non-canonical works. The central hypothesis of the study is that the three text types…

计算与语言 · 计算机科学 2021-04-19 Mahdi Mohseni , Volker Gast , Christoph Redies

Data representation is a fundamental task in machine learning. The representation of data affects the performance of the whole machine learning system. In a long history, the representation of data is done by feature engineering, and…

计算与语言 · 计算机科学 2016-11-21 Siwei Lai

Machine-translated text plays an important role in modern life by smoothing communication from various communities using different languages. However, unnatural translation may lead to misunderstanding, a detector is thus needed to avoid…

计算与语言 · 计算机科学 2019-04-25 Hoang-Quoc Nguyen-Son , Tran Phuong Thao , Seira Hidano , Shinsaku Kiyomoto

Social media data is a valuable resource for research, yet it contains a wide range of non-standard words (NSW). These irregularities hinder the effective operation of NLP tools. Current state-of-the-art methods for the Vietnamese language…

计算与语言 · 计算机科学 2024-07-26 Anh Thi-Hoang Nguyen , Dung Ha Nguyen , Nguyet Thi Nguyen , Khanh Thanh-Duy Ho , Kiet Van Nguyen

In this work we investigate the accuracy of standard and state-of-the-art language identification methods in identifying Albanian in written text documents. A dataset consisting of news articles written in Albanian has been constructed for…

信息检索 · 计算机科学 2019-01-15 Klesti Hoxha , Artur Baxhaku

In this study, we investigate the application of keyword spotting (KWS) in the domain of Hindi speech recognition, utilizing a dataset comprising 40,000 audio samples. With a sampling rate of 44 kHz and an average duration of 1.9 seconds…

声音 · 计算机科学 2026-05-06 Saru Bharti , Pushparaj Mani Pathak

Single document summarization generates summary by extracting the representative sentences from the document. In this paper, we presented a novel technique for summarization of domain-specific text from a single web document that uses…

信息检索 · 计算机科学 2016-11-17 Rushdi Shams , M. M. A. Hashem , Afrina Hossain , Suraiya Rumana Akter , Monika Gope

The Chapter starts with introductory information about quantitative linguistics notions, like rank--frequency dependence, Zipf's law, frequency spectra, etc. Similarities in distributions of words in texts with level occupation in quantum…

数据分析、统计与概率 · 物理学 2024-01-04 Andrij Rovenchak