中文
相关论文

相关论文: The Grammar and Syntax Based Corpus Analysis Tool …

200 篇论文

Stylometry is the study of the unique linguistic styles and writing behaviors of individuals. It belongs to the core task of text categorization like authorship identification, plagiarism detection etc. Though reasonable number of studies…

计算与语言 · 计算机科学 2013-02-26 Tanmoy Chakraborty

Machine translation is evolving quite rapidly in terms of quality. Nowadays, we have several machine translation systems available in the web, which provide reasonable translations. However, these systems are not perfect, and their quality…

计算与语言 · 计算机科学 2015-10-16 Krzysztof Wołk , Krzysztof Marasek , Wojciech Glinkowski

Large language models are demonstrating increasing capabilities, excelling at benchmarks once considered very difficult. As their capabilities grow, there is a need for more challenging evaluations that go beyond surface-level linguistic…

计算与语言 · 计算机科学 2025-10-27 Mojca Brglez , Špela Vintar

Linguistic features remain essential for interpretability and tasks that involve style, structure, and readability, but existing Spanish tools offer limited coverage. We present PUCP-Metrix, an open-source and comprehensive toolkit for…

计算与语言 · 计算机科学 2025-12-05 Javier Alonso Villegas Luis , Marco Antonio Sobrevilla Cabezudo

We present a speech database and a phoneme-level language model of Polish. The database and model are designed for the analysis of prosodic and discourse factors and their impact on acoustic parameters in interaction with predictability…

计算与语言 · 计算机科学 2024-04-19 Zofia Malisz , Jan Foremski , Małgorzata Kul

This study presents a benchmark for evaluating the Visual Word Sense Disambiguation (Visual-WSD) task in Ukrainian. The main goal of the Visual-WSD task is to identify, with minimal contextual information, the most appropriate…

计算机视觉与模式识别 · 计算机科学 2026-03-26 Yurii Laba , Yaryna Mohytych , Ivanna Rohulia , Halyna Kyryleyza , Hanna Dydyk-Meush , Oles Dobosevych , Rostyslav Hryniv

Stylometric analysis of medieval vernacular texts is still a significant challenge: the importance of scribal variation, be it spelling or more substantial, as well as the variants and errors introduced in the tradition, complicate the task…

计算与语言 · 计算机科学 2020-12-08 Jean-Baptiste Camps , Thibault Clérice , Ariane Pinche

Text alignment is crucial to the accuracy of Machine Translation (MT) systems, some NLP tools or any other text processing tasks requiring bilingual data. This research proposes a language independent sentence alignment approach based on…

计算与语言 · 计算机科学 2015-10-01 Krzysztof Wołk , Krzysztof Marasek

The algorithm of the creation texts parallel corpora was presented. The algorithm is based on the use of "key words" in text documents, and on the means of their automated translation. Key words were singled out by means of using Russian…

计算与语言 · 计算机科学 2008-07-03 D. V. Lande , V. V. Zhygalo

Handwritten text generation (HTG) conditioned on writer style has been widely studied for Latin scripts, but remains underexplored for low-resource and non-Latin writing systems, leaving open how well existing models generalise beyond the…

计算机视觉与模式识别 · 计算机科学 2026-05-28 Andrii Ahitoliev , Pavlo Berezin

This paper provides an overview of the morphology and syntax of the Tamil language, focusing on its contemporary usage. The paper also highlights the complexity and richness of Tamil in terms of its morphological and syntactic features,…

计算与语言 · 计算机科学 2024-01-17 Kengatharaiyer Sarveswaran

The emergence of large language models (LLMs) capable of generating realistic texts and images has sparked ethical concerns across various sectors. In response, researchers in academia and industry are actively exploring methods to…

计算与语言 · 计算机科学 2024-05-17 Chidimma Opara

We present our experience in applying distributional semantics (neural word embeddings) to the problem of representing and clustering documents in a bilingual comparable corpus. Our data is a collection of Russian and Ukrainian academic…

计算与语言 · 计算机科学 2016-04-20 Andrey Kutuzov , Mikhail Kopotev , Tatyana Sviridenko , Lyubov Ivanova

We argue that grammatical analysis is a viable alternative to concept spotting for processing spoken input in a practical spoken dialogue system. We discuss the structure of the grammar, and a model for robust parsing which combines…

计算与语言 · 计算机科学 2016-08-31 Gertjan van Noord , Gosse Bouma , Rob Koeling , Mark-Jan Nederhof

We introduce a new Slovak masked language model called SlovakBERT. This is to our best knowledge the first paper discussing Slovak transformers-based language models. We evaluate our model on several NLP tasks and achieve state-of-the-art…

Evaluating writing quality is complex and time-consuming often delaying feedback to learners. While automated writing evaluation tools are effective for English, Korean automated writing evaluation tools face challenges due to their…

计算与语言 · 计算机科学 2025-02-17 Seokho Ahn , Junhyung Park , Ganghee Go , Chulhui Kim , Jiho Jung , Myung Sun Shin , Do-Guk Kim , Young-Duk Seo

While the use of statistical physics methods to analyze large corpora has been useful to unveil many patterns in texts, no comprehensive investigation has been performed investigating the properties of statistical measurements across…

Recent developments in neural language models (LMs) have raised concerns about their potential misuse for automatically spreading misinformation. In light of these concerns, several studies have proposed to detect machine-generated fake…

计算与语言 · 计算机科学 2020-02-21 Tal Schuster , Roei Schuster , Darsh J Shah , Regina Barzilay

Structure or projectional editors are a well-studied concept among researchers and some practitioners. They have the huge advantage of preventing syntax and in some cases type errors, and aid the discovery of syntax by users unfamiliar with…

人机交互 · 计算机科学 2021-07-20 Maryam Hosseinkord , Gurleen Dulai , Narges Osmani , Christopher K. Anand

We analyze the rank-frequency distributions of words in selected English and Polish texts. We compare scaling properties of these distributions in both languages. We also study a few small corpora of Polish literary texts and find that for…

计算与语言 · 计算机科学 2015-11-11 Stanislaw Drozdz , Jaroslaw Kwapien , Adam Orczyk