中文
相关论文

相关论文: Using a Corpus for Teaching Turkish Morphology

200 篇论文

This article presents a hybrid methodology for building a multilingual corpus designed to support the study of emerging concepts in the humanities and social sciences (HSS), illustrated here through the case of ``non-technological…

计算与语言 · 计算机科学 2025-12-09 Revekka Kyriakoglou , Anna Pappa

Tagged corpora play a crucial role in a wide range of Natural Language Processing. The Part of Speech Tagging (POST) is essential in developing tagged corpora. It is time-and-effort-consuming and costly, and therefore, it could be more…

计算与语言 · 计算机科学 2022-02-01 Hossein Hassani

This article investigates the use of Transformation-Based Error-Driven learning for resolving part-of-speech ambiguity in the Greek language. The aim is not only to study the performance, but also to examine its dependence on different…

计算与语言 · 计算机科学 2007-05-23 G. Petasis , G. Paliouras , V. Karkaletsis , C. D. Spyropoulos , I. Androutsopoulos

In this thesis, we developed a comprehensive framework for sentiment analysis that takes its many aspects into account mainly for Turkish. We have also proposed several approaches specific to sentiment analysis in English only. We have…

计算与语言 · 计算机科学 2025-12-02 Cem Rifki Aydin

Tokenization is a pivotal design choice for neural language modeling in morphologically rich languages (MRLs) such as Turkish, where productive agglutination challenges both vocabulary efficiency and morphological fidelity. Prior studies…

计算与语言 · 计算机科学 2026-02-09 Duygu Altinok

Morphological analysis is the study of the formation and structure of words. It plays a crucial role in various tasks in Natural Language Processing (NLP) and Computational Linguistics (CL) such as machine translation and text and speech…

计算与语言 · 计算机科学 2020-05-22 Sina Ahmadi , Hossein Hassani

With the expanding growth of Arabic electronic data on the web, extracting information, which is actually one of the major challenges of the question-answering, is essentially used for building corpus of documents. In fact, building a…

信息检索 · 计算机科学 2018-05-24 Patrice Bellot , Wided Bakari , Mahmoud Neji

Most previous work on the recently developed language-modeling approach to information retrieval focuses on document-specific characteristics, and therefore does not take into account the structure of the surrounding corpus. We propose a…

信息检索 · 计算机科学 2007-05-23 Oren Kurland , Lillian Lee

Sophisticated grammatical error detection/correction tools are available for a small set of languages such as English and Chinese. However, it is not straightforward -- if not impossible -- to adapt them to morphologically rich languages…

计算与语言 · 计算机科学 2024-10-17 Ali Gebeşçe , Gözde Gül Şahin

The developments that language models have provided in fulfilling almost all kinds of tasks have attracted the attention of not only researchers but also the society and have enabled them to become products. There are commercially…

We used Lemur Toolkit, an open source toolkit designed for Information Retrieval (IR) research, for our automated indexing and retrieval experiments on a TREC-like test collection for Turkish. We study and compare three retrieval models…

信息检索 · 计算机科学 2014-05-09 Kutlu Emre Yılmaz , Ahmet Arslan , Ozgur Yilmazel

Urdu is a challenging language because of, first, its Perso-Arabic script and second, its morphological system having inherent grammatical forms and vocabulary of Arabic, Persian and the native languages of South Asia. This paper describes…

计算与语言 · 计算机科学 2022-04-08 Muhammad Humayoun , Harald Hammarström , Aarne Ranta

The word embedding methods have been proven to be very useful in many tasks of NLP (Natural Language Processing). Much has been investigated about word embeddings of English words and phrases, but only little attention has been dedicated to…

计算与语言 · 计算机科学 2016-08-03 Lukáš Svoboda , Tomáš Brychcín

Researchers working in areas such as lexicography, translation studies, and computational linguistics, use a combination of automated and semi-automated tools to analyze the content of text corpora. Keywords, named entities, and events are…

人机交互 · 计算机科学 2022-03-24 Shane Sheehan , Saturnino Luz , Masood Masoodian

In terms of annotation structure, most learner corpora rely on holistic flat label inventories which, even when extensive, do not explicitly separate multiple linguistic dimensions. This makes linguistically deep annotation difficult and…

We illustrate the use of machine learning techniques to analyze, structure, maintain, and evolve a large online corpus of academic literature. An emerging field of research can be identified as part of an existing corpus, permitting the…

信息检索 · 计算机科学 2009-11-10 Paul Ginsparg , Paul Houle , Thorsten Joachims , Jae-Hoon Sul

This paper introduces the first Turkish crossword puzzle generator designed to leverage the capabilities of large language models (LLMs) for educational purposes. In this work, we introduced two specially created datasets: one with over…

计算与语言 · 计算机科学 2024-05-16 Kamyar Zeinalipour , Yusuf Gökberk Keptiğ , Marco Maggini , Leonardo Rigutini , Marco Gori

We introduce Morse, a recurrent encoder-decoder model that produces morphological analyses of each word in a sentence. The encoder turns the relevant information about the word and its context into a fixed size vector representation and the…

计算与语言 · 计算机科学 2019-09-25 Ekin Akyürek , Erenay Dayanık , Deniz Yuret

This paper will present textual corpora for Serbian (and Serbo-Croatian), usable for the training of large language models and publicly available at one of the several notable online repositories. Each corpus will be classified using…

计算与语言 · 计算机科学 2024-05-16 Mihailo Škorić , Nikola Janković

For any deep computational processing of language we need evidences, and one such set of evidences is corpus. This paper describes the development of a text-based corpus for the Bishnupriya Manipuri language. A Corpus is considered as a…

计算与语言 · 计算机科学 2013-12-12 Nayan Jyoti Kalita , Navanath Saharia , Smriti Kumar Sinha