中文
相关论文

相关论文: The Grammar and Syntax Based Corpus Analysis Tool …

200 篇论文

In the era of social networks and rapid misinformation spread, news analysis remains a critical task. Detecting fake news across multiple languages, particularly beyond English, poses significant challenges. Cross-lingual news comparison…

计算与语言 · 计算机科学 2025-10-23 Daryna Dementieva , Evgeniya Sukhodolskaya , Alexander Fraser

We propose a theoretical framework within which information on the vocabulary of a given corpus can be inferred on the basis of statistical information gathered on that corpus. Inferences can be made on the categories of the words in the…

计算与语言 · 计算机科学 2008-10-08 Pascal Vaillant , Richard Nock , Claudia Henry

An evaluation metric is an absolute necessity for measuring the performance of any system and complexity of any data. In this paper, we have discussed how to determine the level of complexity of code-mixed social media texts that are…

计算与语言 · 计算机科学 2017-07-06 Souvick Ghosh , Satanu Ghosh , Dipankar Das

Defining psycholinguistic characteristics in written texts is a task gaining increasing attention from researchers. One of the most widely used tools in the current field is Linguistic Inquiry and Word Count (LIWC) that originally was…

计算与语言 · 计算机科学 2026-01-29 Elina Sigdel , Anastasia Panfilova

As the name suggests, type-logical grammars are a grammar formalism based on logic and type theory. From the prespective of grammar design, type-logical grammars develop the syntactic and semantic aspects of linguistic phenomena…

计算与语言 · 计算机科学 2016-08-29 Richard Moot

TextDescriptives is a Python package for calculating a large variety of metrics from text. It is built on top of spaCy and can be easily integrated into existing workflows. The package has already been used for analysing the linguistic…

计算与语言 · 计算机科学 2023-10-30 Lasse Hansen , Ludvig Renbo Olsen , Kenneth Enevoldsen

Text summarization is an essential task in natural language processing, and researchers have developed various approaches over the years, ranging from rule-based systems to neural networks. However, there is no single model or approach that…

计算与语言 · 计算机科学 2023-08-08 Aleš Žagar , Marko Robnik-Šikonja

Natural language processing tools have become frequently used in social sciences such as economics, political science, and sociology. Many publications apply topic modeling to elicit latent topics in text corpora and their development over…

综合经济学 · 经济学 2024-04-30 W. Benedikt Schmal

This paper refers to the syntactic analysis of phrases in Romanian, as an important process of natural language processing. We will suggest a real-time solution, based on the idea of using some words or groups of words that indicate…

计算与语言 · 计算机科学 2013-01-10 Bogdan Patrut

We use the rank-frequency analysis for the estimation of Kernel Vocabulary size within specific corpora of Ukrainian. The extrapolation of high-rank behaviour is utilized for estimation of the total vocabulary size.

计算与语言 · 计算机科学 2007-05-23 Solomija N. Buk , Andrij A. Rovenchak

Semantic textual similarity (STS) plays a crucial role in many natural language processing tasks. While extensively studied in high-resource languages, STS remains challenging for under-resourced languages such as Slovak. This paper…

计算与语言 · 计算机科学 2026-02-05 Lukas Radosky , Miroslav Blstak , Matej Krajcovic , Ivan Polasek

Grammatic is a tool for grammar definition and manipulation aimed to improve modularity and reuse of grammars and related development artifacts. It is independent from parsing technology and any other details of target system…

编程语言 · 计算机科学 2009-01-19 Andrey Breslav

Spelling error correction is an important problem in natural language processing, as a prerequisite for good performance in downstream tasks as well as an important feature in user-facing applications. For texts in Polish language, there…

计算与语言 · 计算机科学 2019-05-28 Szymon Rutkowski

Forensic scientists often need to identify an unknown speaker or writer in cases such as ransom calls, covert recordings, alleged suicide notes, or anonymous online communications, among many others. Speaker recognition in the speech domain…

计算与语言 · 计算机科学 2025-12-19 Cristina Aggazzotti , Elizabeth Allyn Smith

This research introduces a novel psychometric method for analyzing textual data using large language models. By leveraging contextual embeddings to create contextual scores, we transform textual data into response data suitable for…

计算与语言 · 计算机科学 2025-09-12 Jinsong Chen

Grammar checking is the task of detection and correction of grammatical errors in the text. English is the dominating language in the field of science and technology. Therefore, the non-native English speakers must be able to use correct…

计算与语言 · 计算机科学 2018-04-03 Madhvi Soni , Jitendra Singh Thakur

We present a cross-lingual summarisation corpus with long documents in a source language associated with multi-sentence summaries in a target language. The corpus covers twelve language pairs and directions for four European languages,…

计算与语言 · 计算机科学 2022-02-22 Laura Perez-Beltrachini , Mirella Lapata

Analyzing textual data is a very challenging task because of the huge volume of data generated daily. Fundamental issues in text analysis include the lack of structure in document datasets, the need for various preprocessing steps %(e.g.,…

数据库 · 计算机科学 2016-12-20 Ciprian-Octavian Truică , Jérôme Darmont , Julien Velcin

High-quality automated poetry generation systems are currently only available for a small subset of languages. We introduce a new model for generating poetry in Czech language, based on fine-tuning a pre-trained Large Language Model. We…

计算与语言 · 计算机科学 2024-07-19 Michal Chudoba , Rudolf Rosa

The rapid advancement of large language models (LLMs) has made it increasingly difficult to distinguish between text written by humans and machines. Addressing this, we propose a novel method for generating watermarks that strategically…

计算与语言 · 计算机科学 2024-05-15 Georg Niess , Roman Kern