English
Related papers

Related papers: The Grammar and Syntax Based Corpus Analysis Tool …

200 papers

Norms, which are culturally accepted guidelines for behaviours, can be integrated into conversational models to generate utterances that are appropriate for the socio-cultural context. Existing methods for norm recognition tend to focus…

Computation and Language · Computer Science 2023-05-29 Farhad Moghimifar , Shilin Qu , Tongtong Wu , Yuan-Fang Li , Gholamreza Haffari

Mathematical documents written in LaTeX often contain ambiguities. We can resolve some of them via semantic markup using, e.g., sTeX, which also has other potential benefits, such as interoperability with computer algebra systems, proof…

Computation and Language · Computer Science 2024-08-12 Luka Vrečar , Joe Wells , Fairouz Kamareddine

Language identification is an important Natural Language Processing task. It has been thoroughly researched in the literature. However, some issues are still open. This work addresses the identification of the related low-resource languages…

Computation and Language · Computer Science 2022-03-10 Olha Dovbnia , Anna Wróblewska

Text generation systems are ubiquitous in natural language processing applications. However, evaluation of these systems remains a challenge, especially in multilingual settings. In this paper, we propose L'AMBRE -- a metric to evaluate the…

Computation and Language · Computer Science 2021-09-13 Adithya Pratapa , Antonios Anastasopoulos , Shruti Rijhwani , Aditi Chaudhary , David R. Mortensen , Graham Neubig , Yulia Tsvetkov

The era of large language models (LLM) raises questions not only about how to train models, but also about how to evaluate them. Despite numerous existing benchmarks, insufficient attention is often given to creating assessments that test…

We are concerned with the syntactic annotation of unrestricted text. We combine a rule-based analysis with subsequent exploitation of empirical data. The rule-based surface syntactic analyser leaves some amount of ambiguity in the output…

cmp-lg · Computer Science 2008-02-03 Pasi Tapanainen , Timo Järvinen

Large language models (LLMs) can be used to generate smaller, more refined datasets via few-shot prompting for benchmarking, fine-tuning or other use cases. However, understanding and evaluating these datasets is difficult, and the failure…

Computation and Language · Computer Science 2023-09-29 Emily Reif , Minsuk Kahng , Savvas Petridis

This paper presents the first comprehensive study on automatic readability assessment of Turkish texts. We combine state-of-the-art neural network models with linguistic features at lexical, morphological, syntactic and discourse levels to…

Computation and Language · Computer Science 2025-09-05 Ahmet Yavuz Uluslu , Gerold Schneider

In recent years, sentiment analysis has gained increasing significance, prompting researchers to explore datasets in various languages, including Turkish. However, the limited availability of Turkish datasets has led to their multifaceted…

Computation and Language · Computer Science 2024-12-10 Şevval Çakıcı , Dilara Karaduman , Mehmet Akif Çırlan , Ali Hürriyetoğlu

Texts exhibit considerable stylistic variation. This paper reports an experiment where a corpus of documents (N= 75 000) is analyzed using various simple stylistic metrics. A subset (n = 1000) of the corpus has been previously assessed to…

cmp-lg · Computer Science 2008-02-03 Jussi Karlgren

Text summarization is the task of shortening a larger body of text into a concise version while retaining its essential meaning and key information. While summarization has been significantly explored in English and other high-resource…

Computation and Language · Computer Science 2025-08-15 Václav Tran , Jakub Šmíd , Jiří Martínek , Ladislav Lenc , Pavel Král

We present Charles University submissions to the WMT22 General Translation Shared Task on Czech-Ukrainian and Ukrainian-Czech machine translation. We present two constrained submissions based on block back-translation and tagged…

Computation and Language · Computer Science 2022-12-02 Martin Popel , Jindřich Libovický , Jindřich Helcl

The text editing tasks, including sentence fusion, sentence splitting and rephrasing, text simplification, and Grammatical Error Correction (GEC), share a common trait of dealing with highly similar input and output sequences. This area of…

Computation and Language · Computer Science 2023-09-21 Bohdan Didenko , Andrii Sameliuk

Existing corpora for intrinsic evaluation are not targeted towards tasks in informal domains such as Twitter or news comment forums. We want to test whether a representation of informal words fulfills the promise of eliding explicit text…

Computation and Language · Computer Science 2016-06-28 Naomi Saphra , Adam Lopez

Text analysis of social media for sentiment, topic analysis, and other analysis depends initially on the selection of keywords and phrases that will be used to create the research corpora. However, keywords that researchers choose may occur…

Computation and Language · Computer Science 2022-04-21 Philip Feldman , Aaron Dant , James R. Foulds , Shemei Pan

Leading large language models have demonstrated impressive capabilities in reasoning-intensive tasks, such as standardized educational testing. However, they often require extensive training in low-resource settings with inaccessible…

Computation and Language · Computer Science 2025-03-19 Mykyta Syromiatnikov , Victoria Ruvinskaya , Nataliia Komleva

Recent advances in text mining and natural language processing technology have enabled researchers to detect an authors identity or demographic characteristics, such as age and gender, in several text genres by automatically analysing the…

Cryptography and Security · Computer Science 2022-11-30 Claudia Peersman , Matthew Edwards , Emma Williams , Awais Rashid

By representing a text by a set of words and their co-occurrences, one obtains a word-adjacency network being a reduced representation of a given language sample. In this paper, the possibility of using network representation to extract…

Computation and Language · Computer Science 2019-01-18 Tomasz Stanisz , Jarosław Kwapień , Stanisław Drożdż

Small Language Models (SLMs) have become increasingly important due to their efficiency and performance to perform various language tasks with minimal computational resources, making them ideal for various settings including on-device,…

A pronoun resolution system which requires limited syntactic knowledge to identify the antecedents of personal and reflexive pronouns in Turkish is presented. As in its counterparts for languages like English, Spanish and French, the core…

Computation and Language · Computer Science 2015-04-21 Dilek Küçük , Meltem Turhan Yöndem
‹ Prev 1 8 9 10 Next ›