中文
相关论文

相关论文: Automatic Segmentation of Manipuri (Meiteilon) Wor…

200 篇论文

This paper introduces a new score-informed method for the segmentation of jingju a cappella singing phrase into syllables. The proposed method estimates the most likely sequence of syllable boundaries given the estimated syllable onset…

声音 · 计算机科学 2017-07-13 Jordi Pons , Rong Gong , Xavier Serra

Based on the sense definition of words available in the Bengali WordNet, an attempt is made to classify the Bengali sentences automatically into different groups in accordance with their underlying senses. The input sentences are collected…

计算与语言 · 计算机科学 2015-08-07 Alok Ranjan Pal , Diganta Saha , Niladri Sekhar Dash

Spelling errors are introduced in text either during typing, or when the user does not know the correct phoneme or grapheme. If a language contains complex words like sandhi where two or more morphemes join based on some rules, spell…

计算与语言 · 计算机科学 2016-11-28 A N Akshatha , Chandana G Upadhyaya , Rajashekara S Murthy

Recent years the task of incomplete utterance rewriting has raised a large attention. Previous works usually shape it as a machine translation task and employ sequence to sequence based architecture with copy mechanism. In this paper, we…

计算与语言 · 计算机科学 2020-09-29 Qian Liu , Bei Chen , Jian-Guang Lou , Bin Zhou , Dongmei Zhang

In Neural Machine Translation (NMT) the usage of subwords and characters as source and target units offers a simple and flexible solution for translation of rare and unseen words. However, selecting the optimal subword segmentation involves…

计算与语言 · 计算机科学 2019-10-29 Tejas Srinivasan , Ramon Sanabria , Florian Metze

Handwriting recognition remains challenging for some of the most spoken languages, like Bangla, due to the complexity of line and word segmentation brought by the curvilinear nature of writing and lack of quality datasets. This paper solves…

计算机视觉与模式识别 · 计算机科学 2023-09-27 Sheikh Mohammad Jubaer , Nazifa Tabassum , Md. Ataur Rahman , Mohammad Khairul Islam

Neural machine translation (NMT) models are typically trained with fixed-size input and output vocabularies, which creates an important bottleneck on their accuracy and generalization capability. As a solution, various studies proposed…

计算与语言 · 计算机科学 2018-05-08 Duygu Ataman , Marcello Federico

Indexed languages are a classical notion in formal language theory. As the language equivalent of second-order pushdown automata, they have received considerable attention in higher-order model checking. Unfortunately, counting properties…

形式语言与自动机理论 · 计算机科学 2024-05-14 Laura Ciobanu , Georg Zetzsche

There is an abundance of digitised texts available in Sanskrit. However, the word segmentation task in such texts are challenging due to the issue of 'Sandhi'. In Sandhi, words in a sentence often fuse together to form a single chunk of…

计算与语言 · 计算机科学 2018-02-20 Vikas Reddy , Amrith Krishna , Vishnu Dutt Sharma , Prateek Gupta , Vineeth M R , Pawan Goyal

Neural models with minimal feature engineering have achieved competitive performance against traditional methods for the task of Chinese word segmentation. However, both training and working procedures of the current neural models are…

计算与语言 · 计算机科学 2017-04-25 Deng Cai , Hai Zhao , Zhisong Zhang , Yuan Xin , Yongjian Wu , Feiyue Huang

A novel approach for speech segmentation is proposed, based on Multilevel Hybrid (mean/min) Filters (MHF) with the following features: An accurate transition location. Good performance in noisy environments (gaussian and impulsive noise).…

音频与语音处理 · 电气工程与系统科学 2022-03-04 Marcos Faundez-Zanuy , Francesc Vallverdu-Bayes

This paper describes word {segmentation} granularity in Korean language processing. From a word separated by blank space, which is termed an eojeol, to a sequence of morphemes in Korean, there are multiple possible levels of word…

计算与语言 · 计算机科学 2023-09-08 Jungyeul Park , Mija Kim

Dividing oral histories into topically coherent segments can make them more accessible online. People regularly make judgments about where coherent segments can be extracted from oral histories. But making these judgments can be taxing, so…

计算与语言 · 计算机科学 2015-09-30 Ryan Shaw

Tibetan is a low-resource language. In order to alleviate the shortage of parallel corpus between Tibetan and Chinese, this paper uses two monolingual corpora and a small number of seed dictionaries to learn the semi-supervised method with…

计算与语言 · 计算机科学 2021-10-05 Enshuai Hou , Jie zhu

In recent decades, Speech interactive systems gained increasing importance. To develop Dictation System like Dragon for Indian languages it is most important to adapt the system to a speaker with minimum training. In this paper we focus on…

计算与语言 · 计算机科学 2010-01-14 N. Kalyani , Dr K. V. N. Sunitha

Despite being one of the most spoken languages in the world ($6^{th}$ based on population), research regarding Bengali handwritten grapheme (smallest functional unit of a writing system) classification has not been explored widely compared…

计算机视觉与模式识别 · 计算机科学 2021-11-17 Tarun Roy , Hasib Hasan , Kowsar Hossain , Masuma Akter Rumi

We propose a new approach to the Chinese word segmentation problem that considers the sentence as an undirected graph, whose nodes are the characters. One can use various techniques to compute the edge weights that measure the connection…

计算与语言 · 计算机科学 2018-04-06 Yuanhao Liu , Sheng Yu

Developing benchmark datasets for low-resource languages poses significant challenges, primarily due to the limited availability of native linguistic experts and the substantial time and cost involved in annotation. Given these challenges,…

计算与语言 · 计算机科学 2025-10-28 Rahul Ranjan , Mahendra Kumar Gurve , Anuj , Nitin , Yamuna Prasad

Sandhi means to join two or more words to coin new word. Sandhi literally means `putting together' or combining (of sounds), It denotes all combinatory sound-changes effected (spontaneously) for ease of pronunciation. Sandhi-vicheda…

计算与语言 · 计算机科学 2009-09-15 Priyanka Gupta , Vishal Goyal

In neural machine translation (NMT), it is has become standard to translate using subword units to allow for an open vocabulary and improve accuracy on infrequent words. Byte-pair encoding (BPE) and its variants are the predominant approach…

计算与语言 · 计算机科学 2018-10-23 Elizabeth Salesky , Andrew Runge , Alex Coda , Jan Niehues , Graham Neubig