English
Related papers

Related papers: Keyboards for inputting Japanese language -A study…

200 papers

The old Bulgarian keyboard standard BDS 5237-78 was developed for use mostly on typewriter machines. The wide distribution of the computers forced the update of this standard. On one hand this is because of the need to support symbols such…

Human-Computer Interaction · Computer Science 2009-05-06 Anton Zinoviev

A keyboard has many function keys and each function key can have multiple functions when used with control, shift and alt keys, it is difficult for a user to remember the functionality of the function keys. We need a mechanism to indicate…

Human-Computer Interaction · Computer Science 2013-10-14 Umakant Mishra

The majority of online content is written in languages other than English, and is most commonly encoded in UTF-8, the world's dominant Unicode character encoding. Traditional compression algorithms typically operate on individual bytes.…

Information Theory · Computer Science 2017-01-17 Adam Gleave , Christian Steinruecken

Intelligent Input Methods (IM) are essential for making text entries in many East Asian scripts, but their application to other languages has not been fully explored. This paper discusses how such tools can contribute to the development of…

Computation and Language · Computer Science 2007-05-23 Mike Tian-Jian Jiang , Deng Liu , Meng-Juei Hsieh , Wen-Lien Hsu

Keyboard, although a popular medium, is not very convenient as it requires a certain amount of skill for effective usage. A mouse on the other hand requires a good hand-eye co-ordination. Also current computer interfaces also assume a…

Human-Computer Interaction · Computer Science 2014-01-16 Kamlesh Sharma , Dr. T. V Prasad

Chinese input methods are used to convert pinyin sequence or other Latin encoding systems into Chinese character sentences. For more effective pinyin-to-character conversion, typical Input Method Engines (IMEs) rely on a predefined…

Computation and Language · Computer Science 2017-12-13 Xihu Zhang , Chu Wei , Hai Zhao

Language expression is known to be dependent on attributes intrinsic to the author. To date, however, little attention has been devoted to the effect of interfaces used to articulate language on its expression. Here we study a large corpus…

Human-Computer Interaction · Computer Science 2015-05-04 Dan Pelleg , Elad Yom-Tov , Evgeniy Gabrilovich

Text-based passwords continue to be the prime form of authentication to computer systems. Today, they are increasingly created and used with mobile text entry methods, such as touchscreens and mobile keyboards, in addition to traditional…

Cryptography and Security · Computer Science 2014-03-11 Yulong Yang , Janne Lindqvist , Antti Oulasvirta

In this age of information technology, information access in a convenient manner has gained importance. Since speech is a primary mode of communication among human beings, it is natural for people to expect to be able to carry out spoken…

Computation and Language · Computer Science 2013-05-14 Neema Mishra , Urmila Shrawankar , V M Thakare

This paper investigates the effect of tokenizers on the downstream performance of pretrained language models (PLMs) in scriptio continua languages where no explicit spaces exist between words, using Japanese as a case study. The tokenizer…

Computation and Language · Computer Science 2023-06-19 Takuro Fujii , Koki Shibata , Atsuki Yamaguchi , Terufumi Morishita , Yasuhiro Sogawa

We introduce a Japanese Morphology dataset, J-UniMorph, developed based on the UniMorph feature schema. This dataset addresses the unique and rich verb forms characteristic of the language's agglutinative nature. J-UniMorph distinguishes…

Computation and Language · Computer Science 2024-02-23 Kosuke Matsuzaki , Masaya Taniguchi , Kentaro Inui , Keisuke Sakaguchi

Korean is a morphologically rich language with a featural writing system in which each character is systematically composed of subcharacter units known as Jamo. These subcharacters not only determine the visual structure of Korean but also…

Computation and Language · Computer Science 2026-04-15 SungHo Kim , Juhyeong Park , Eda Atalay , SangKeun Lee

This paper shows the necessity of distinguishing different referential uses of noun phrases in machine translation. We argue that differentiating between the generic, referential and ascriptive uses of noun phrases is the minimum necessary…

cmp-lg · Computer Science 2008-02-03 Francis Bond , Kentaro Ogura , Tsukasa Kawaoka

While translating between East Asian languages, many works have discovered clear advantages of using characters as the translation unit. Unfortunately, traditional recurrent neural machine translation systems hinder the practical usage of…

Computation and Language · Computer Science 2020-12-17 Thi-Vinh Ngo , Thanh-Le Ha , Phuong-Thai Nguyen , Le-Minh Nguyen

Chinese calligraphy is the writing of Chinese characters as an art form performed with brushes so Chinese characters are rich of shapes and details. Recent studies show that Chinese characters can be generated through image-to-image…

Computer Vision and Pattern Recognition · Computer Science 2020-05-27 Shan-Jean Wu , Chih-Yuan Yang , Jane Yung-jen Hsu

Almost all existing machine translation models are built on top of character-based vocabularies: characters, subwords or words. Rare characters from noisy text or character-rich languages such as Japanese and Chinese however can…

Computation and Language · Computer Science 2019-12-09 Changhan Wang , Kyunghyun Cho , Jiatao Gu

Open Japanese large language models (LLMs) have been trained on the Japanese portions of corpora such as CC-100, mC4, and OSCAR. However, these corpora were not created for the quality of Japanese texts. This study builds a large Japanese…

Computation and Language · Computer Science 2024-04-30 Naoaki Okazaki , Kakeru Hattori , Hirai Shota , Hiroki Iida , Masanari Ohi , Kazuki Fujii , Taishi Nakamura , Mengsay Loem , Rio Yokota , Sakae Mizuki

The national languages of Senegal, like those of West Africa country in general, are written with two alphabets : the Latin alphabet that draws its strength from official decreesm and the completed Arabic script (Ajami), widespread and well…

Computation and Language · Computer Science 2020-05-07 El hadji M. Fall , El hadji M. Nguer , Bao Diop Sokhna , Mouhamadou Khoule , Mathieu Mangeot , Mame T. Cisse

Previous approaches to training syntax-based sentiment classification models required phrase-level annotated corpora, which are not readily available in many languages other than English. Thus, we propose the use of tree-structured Long…

Computation and Language · Computer Science 2018-10-02 Ryosuke Miyazaki , Mamoru Komachi

This paper describes two new bunsetsu identification methods using supervised learning. Since Japanese syntactic analysis is usually done after bunsetsu identification, bunsetsu identification is important for analyzing Japanese sentences.…

Computation and Language · Computer Science 2007-05-23 Masaki Murata , Kiyotaka Uchimoto , Qing Ma , Hitoshi Isahara