English
Related papers

Related papers: Developing an Informal-Formal Persian Corpus

200 papers

The rapid advancement of language models has demonstrated the potential of artificial intelligence in the healthcare industry. However, small language models struggle with specialized domains in low-resource languages like Persian. While…

Computation and Language · Computer Science 2025-11-18 Mehrdad Ghassabi , Pedram Rostami , Hamidreza Baradaran Kashani , Amirhossein Poursina , Zahra Kazemi , Milad Tavakoli

The present paper investigated automatic melody construction for Persian lyrics as an input. It was assumed that there is a phonological correlation between the lyric syllables and the melody in a song. A seq2seq neural network was…

Sound · Computer Science 2024-10-29 Farshad Jafari , Farzad Didehvar , Amin Gheibi

In this work, we employ a semi-automatic method based on back translation to generate a sentential paraphrase corpus for the Armenian language. The initial collection of sentences is translated from Armenian to English and back twice,…

Computation and Language · Computer Science 2020-09-29 Arthur Malajyan , Karen Avetisyan , Tsolak Ghukasyan

In general, speech processing models consist of a language model along with an acoustic model. Regardless of the language model's complexity and variants, three critical pre-processing steps are needed in language models: cleaning,…

Audio and Speech Processing · Electrical Eng. & Systems 2021-12-16 Romina Oji , Seyedeh Fatemeh Razavi , Sajjad Abdi Dehsorkh , Alireza Hariri , Hadi Asheri , Reshad Hosseini

The democratization of AI is currently hindered by the immense computational costs required to train Large Language Models (LLMs) for low-resource languages. This paper presents Persian-Phi, a 3.8B parameter model that challenges the…

Computation and Language · Computer Science 2025-12-09 Amir Mohammad Akhlaghi , Amirhossein Shabani , Mostafa Abdolmaleki , Saeed Reza Kheradpisheh

Displaying a document in Middle Eastern languages requires contextual analysis due to different presentational forms for each character of the alphabet. The words of the document will be formed by the joining of the correct positional…

Computation and Language · Computer Science 2015-09-15 Kazem Taghva

This paper presents the challenges in creating and managing large parallel corpora of 12 major Indian languages (which is soon to be extended to 23 languages) as part of a major consortium project funded by the Department of Information…

Computation and Language · Computer Science 2021-12-06 Ritesh Kumar , Shiv Bhusan Kaushik , Pinkey Nainwani , Girish Nath Jha

This paper introduces the Persian Abstract Meaning Representation (AMR) guidelines, a detailed guide for annotating Persian sentences with AMR, focusing on the necessary adaptations to fit Persian's unique syntactic structures. We discuss…

Computation and Language · Computer Science 2025-04-22 Reza Takhshid , Tara Azin , Razieh Shojaei , Mohammad Bahrani

Spelling correction is a remarkable challenge in the field of natural language processing. The objective of spelling correction tasks is to recognize and rectify spelling errors automatically. The development of applications that can…

Computation and Language · Computer Science 2024-05-07 Mohammad Dehghani , Heshaam Faili

We present sentence aligned parallel corpora across 10 Indian Languages - Hindi, Telugu, Tamil, Malayalam, Gujarati, Urdu, Bengali, Oriya, Marathi, Punjabi, and English - many of which are categorized as low resource. The corpora are…

Computation and Language · Computer Science 2020-07-16 Shashank Siripragada , Jerin Philip , Vinay P. Namboodiri , C V Jawahar

Despite the progress made in recent years in addressing natural language understanding (NLU) challenges, the majority of this progress remains to be concentrated on resource-rich languages like English. This work focuses on Persian…

India's linguistic landscape is one of the most diverse in the world, comprising over 120 major languages and approximately 1,600 additional languages, with 22 officially recognized as scheduled languages in the Indian Constitution. Despite…

We present PashtoCorp, a 1.25-billion-word corpus for Pashto, a language spoken by 60 million people that remains severely underrepresented in NLP. The corpus is assembled from 39 sources spanning seven HuggingFace datasets and 32…

Computation and Language · Computer Science 2026-03-18 Hanif Rahman

The patterns in which the syntax of different languages converges and diverges are often used to inform work on cross-lingual transfer. Nevertheless, little empirical work has been done on quantifying the prevalence of different syntactic…

Computation and Language · Computer Science 2020-07-14 Dmitry Nikolaev , Ofir Arviv , Taelin Karidi , Neta Kenneth , Veronika Mitnik , Lilja Maria Saeboe , Omri Abend

Paraphrasing is often performed with less concern for controlled style conversion. Especially for questions and commands, style-variant paraphrasing can be crucial in tone and manner, which also matters with industrial applications such as…

Computation and Language · Computer Science 2022-04-29 Won Ik Cho , Sangwhan Moon , Jong In Kim , Seok Min Kim , Nam Soo Kim

Handwriting analysis is still an important application in machine learning. A basic requirement for any handwriting recognition application is the availability of comprehensive datasets. Standard labelled datasets play a significant role in…

Computer Vision and Pattern Recognition · Computer Science 2024-06-04 Pourya Jafarzadeh , Padideh Choobdar , Vahid Mohammadi Safarzadeh

This paper describes a web-based corpus of global language use with a focus on how this corpus can be used for data-driven language mapping. First, the corpus provides a representation of where national varieties of major languages are used…

Computation and Language · Computer Science 2020-04-03 Jonathan Dunn

Large language models demonstrate remarkable proficiency in various linguistic tasks and have extensive knowledge across various domains. Although they perform best in English, their ability in other languages is notable too. In contrast,…

Computation and Language · Computer Science 2024-01-15 Pedram Rostami , Ali Salemi , Mohammad Javad Dousti

We release a corpus of 43 million atomic edits across 8 languages. These edits are mined from Wikipedia edit history and consist of instances in which a human editor has inserted a single contiguous phrase into, or deleted a single…

Computation and Language · Computer Science 2018-08-29 Manaal Faruqui , Ellie Pavlick , Ian Tenney , Dipanjan Das

Stylistic variation in text needs to be studied with different aspects including the writer's personal traits, interpersonal relations, rhetoric, and more. Despite recent attempts on computational modeling of the variation, the lack of…

Computation and Language · Computer Science 2019-09-04 Dongyeop Kang , Varun Gangal , Eduard Hovy
‹ Prev 1 3 4 5 6 7 10 Next ›