English
Related papers

Related papers: J-UniMorph: Japanese Morphological Annotation thro…

200 papers

This paper presents our segmentation system developed for the MLP 2017 shared tasks on cross-lingual word segmentation and morpheme segmentation. We model both word and morpheme segmentation as character-level sequence labelling tasks. The…

Computation and Language · Computer Science 2017-09-13 Yan Shao

This study presents a Long Short-Term Memory (LSTM) neural network approach to Japanese word segmentation (JWS). Previous studies on Chinese word segmentation (CWS) succeeded in using recurrent neural networks such as LSTM and gated…

Computation and Language · Computer Science 2018-09-28 Yoshiaki Kitagawa , Mamoru Komachi

We propose a novel approach to learn word embeddings based on an extended version of the distributional hypothesis. Our model derives word embedding vectors using the etymological composition of words, rather than the context in which they…

Computation and Language · Computer Science 2017-12-13 Seunghyun Yoon , Pablo Estrada , Kyomin Jung

Recent research argues that exact recursive numeral systems optimize communicative efficiency by balancing a tradeoff between the size of the numeral lexicon and the average morphosyntactic complexity (roughly length in morphemes) of…

We propose a novel phrase break prediction method that combines implicit features extracted from a pre-trained large language model, a.k.a BERT, and explicit features extracted from BiLSTM with linguistic features. In conventional BiLSTM…

Audio and Speech Processing · Electrical Eng. & Systems 2021-04-27 Kosuke Futamata , Byeongseon Park , Ryuichi Yamamoto , Kentaro Tachibana

We present an algorithm that takes an unannotated corpus as its input, and returns a ranked list of probable morphologically related pairs as its output. The algorithm tries to discover morphologically related pairs by looking for pairs…

Computation and Language · Computer Science 2007-05-23 Marco Baroni , Johannes Matiasek , Harald Trost

Previous approaches to training syntax-based sentiment classification models required phrase-level annotated corpora, which are not readily available in many languages other than English. Thus, we propose the use of tree-structured Long…

Computation and Language · Computer Science 2018-10-02 Ryosuke Miyazaki , Mamoru Komachi

We introduce Multi-SimLex, a large-scale lexical resource and evaluation benchmark covering datasets for 12 typologically diverse languages, including major languages (e.g., Mandarin Chinese, Spanish, Russian) as well as less-resourced ones…

When translating Japanese nouns into English, we face the problem of articles and numbers which the Japanese language does not have, but which are necessary for the English composition. To solve this difficult problem we classified the…

cmp-lg · Computer Science 2008-02-03 Masaki Murata , Makoto Nagao

The literature on word-representable graphs is quite rich, and a number of variations of the original definition have been proposed over the years. We are initiating a systematic study of such variations based on formal languages. In our…

Discrete Mathematics · Computer Science 2024-11-06 Zhidan Feng , Henning Fernau , Pamela Fleischmann , Kevin Mann , Silas Cato Sacher

The Sejong dictionary dataset offers a valuable resource, providing extensive coverage of morphology, syntax, and semantic representation. This dataset can be utilized to explore linguistic information in greater depth. The labeled…

Computation and Language · Computer Science 2025-04-04 Seohyun Song , Eunkyul Leah Jo , Yige Chen , Jeen-Pyo Hong , Kyuwon Kim , Jin Wee , Miyoung Kang , KyungTae Lim , Jungyeul Park , Chulwoo Park

We present a free Japanese-French parallel corpus. It includes 15M aligned segments and is obtained by compiling and filtering several existing resources. In this paper, we describe the existing resources, their quantity and quality, the…

Computation and Language · Computer Science 2022-08-30 Raoul Blin , Fabien Cromières

The study of linguistic typology is rooted in the implications we find between linguistic features, such as the fact that languages with object-verb word ordering tend to have post-positions. Uncovering such implications typically amounts…

Computation and Language · Computer Science 2019-06-19 Johannes Bjerva , Yova Kementchedjhieva , Ryan Cotterell , Isabelle Augenstein

Humour, a fundamental aspect of human communication, manifests itself in various styles that significantly impact social interactions and mental health. Recognising different humour styles poses challenges due to the lack of established…

Computation and Language · Computer Science 2024-10-18 Mary Ogbuka Kenneth , Foaad Khosmood , Abbas Edalat

Natural language processing models have attracted much interest in the deep learning community. This branch of study is composed of some applications such as machine translation, sentiment analysis, named entity recognition, question and…

Computation and Language · Computer Science 2020-07-22 Flávio Santos , Hendrik Macedo , Thiago Bispo , Cleber Zanchettin

This study explores language change and variation in Toki Pona, a constructed language with approximately 120 core words. Taking a computational and corpus-based approach, the study examines features including fluid word classes and…

Computation and Language · Computer Science 2025-08-15 Daniel Huang , Hyoun-A Joo

We present the first resource focusing on the verbal inflectional morphology of San Juan Quiahije Chatino, a tonal mesoamerican language spoken in Mexico. We provide a collection of complete inflection tables of 198 lemmata, with…

Computation and Language · Computer Science 2020-04-07 Hilaria Cruz , Gregory Stump , Antonios Anastasopoulos

In this paper, we present a kernel-based learning approach for the 2018 Complex Word Identification (CWI) Shared Task. Our approach is based on combining multiple low-level features, such as character n-grams, with high-level semantic…

Computation and Language · Computer Science 2018-05-23 Andrei M. Butnaru , Radu Tudor Ionescu

Handwritten Mathematical Expression Recognition (HMER) remains a persistent challenge in Optical Character Recognition (OCR) due to the inherent freedom of symbol layouts and variability in handwriting styles. Prior methods have faced…

Computer Vision and Pattern Recognition · Computer Science 2025-10-28 Yu Li , Jin Jiang , Jianhua Zhu , Shuai Peng , Baole Wei , Yuxuan Zhou , Liangcai Gao

Open Japanese large language models (LLMs) have been trained on the Japanese portions of corpora such as CC-100, mC4, and OSCAR. However, these corpora were not created for the quality of Japanese texts. This study builds a large Japanese…

Computation and Language · Computer Science 2024-04-30 Naoaki Okazaki , Kakeru Hattori , Hirai Shota , Hiroki Iida , Masanari Ohi , Kazuki Fujii , Taishi Nakamura , Mengsay Loem , Rio Yokota , Sakae Mizuki
‹ Prev 1 8 9 10 Next ›