English
Related papers

Related papers: PersianLLaMA: Towards Building First Persian Large…

200 papers

This research introduces the first large-scale, well-balanced Persian social media text classification dataset, specifically designed to address the lack of comprehensive resources in this domain. The dataset comprises 36,000 posts across…

Computation and Language · Computer Science 2026-05-26 Isun Chehreh , Ebrahim Ansari

The democratization of AI is currently hindered by the immense computational costs required to train Large Language Models (LLMs) for low-resource languages. This paper presents Persian-Phi, a 3.8B parameter model that challenges the…

Computation and Language · Computer Science 2025-12-09 Amir Mohammad Akhlaghi , Amirhossein Shabani , Mostafa Abdolmaleki , Saeed Reza Kheradpisheh

Multilingual Large Language Models (LLMs) often provide suboptimal performance on low-resource languages like Urdu. This paper introduces UrduLLaMA 1.0, a model derived from the open-source Llama-3.1-8B-Instruct architecture and continually…

Computation and Language · Computer Science 2025-02-25 Layba Fiaz , Munief Hassan Tahir , Sana Shams , Sarmad Hussain

Large Language Models (LLMs), trained on extensive datasets using advanced deep learning architectures, have demonstrated remarkable performance across a wide range of language tasks, becoming a cornerstone of modern AI technologies.…

Large language models (LLMs) struggle to navigate culturally specific communication norms, limiting their effectiveness in global contexts. We focus on Persian taarof, a social norm in Iranian interactions, which is a sophisticated system…

Computation and Language · Computer Science 2025-09-03 Nikta Gohari Sadr , Sahar Heidariasl , Karine Megerdoomian , Laleh Seyyed-Kalantari , Ali Emami

The lack of a suitable tool for the analysis of conversational texts in the Persian language has made various analyses of these texts, including Sentiment Analysis, difficult. In this research, we tried to make the understanding of these…

Computation and Language · Computer Science 2024-09-05 Mohsen Khazeni , Mohammad Heydari , Amir Albadvi

Persian remains substantially underrepresented in open speech-text resources, limiting progress in multi-speaker text-to-speech (TTS), speech-language modelling, and low-resource speech processing. We introduce ParsVoice, the largest…

Sound · Computer Science 2026-05-27 Mohammad Javad Ranjbar Kalahroodi , Heshaam Faili , Azadeh Shakery

Language modeling has witnessed remarkable advancements in recent years, with Large Language Models (LLMs) like ChatGPT setting unparalleled benchmarks in human-like text generation. However, a prevailing limitation is the…

Computation and Language · Computer Science 2023-11-13 Abhinand Balachandran

The Iranian Persian language has two varieties: standard and colloquial. Most natural language processing tools for Persian assume that the text is in standard form: this assumption is wrong in many real applications especially web content.…

Computation and Language · Computer Science 2020-12-11 Mohammad Sadegh Rasooli , Farzane Bakhtyari , Fatemeh Shafiei , Mahsa Ravanbakhsh , Chris Callison-Burch

Spelling correction is a remarkable challenge in the field of natural language processing. The objective of spelling correction tasks is to recognize and rectify spelling errors automatically. The development of applications that can…

Computation and Language · Computer Science 2024-05-07 Mohammad Dehghani , Heshaam Faili

The study of historical languages presents unique challenges due to their complex orthographic systems, fragmentary textual evidence, and the absence of standardized digital representations of text in those languages. Tackling these…

Computation and Language · Computer Science 2025-03-25 Farhan Farsi , Parnian Fazel , Sepand Haghighi , Sadra Sabouri , Farzaneh Goshtasb , Nadia Hajipour , Ehsaneddin Asgari , Hossein Sameti

Language models serve as a cornerstone in natural language processing (NLP), utilizing mathematical methods to generalize language laws and knowledge for prediction and generation. Over extensive research spanning decades, language modeling…

Computation and Language · Computer Science 2024-09-24 Zichong Wang , Zhibo Chu , Thang Viet Doan , Shiwen Ni , Min Yang , Wenbin Zhang

The advancements in the Large Language Model (LLM) have helped in solving several problems related to language processing. Most of the researches have focused on the English language only, because of its popularity and abundance on the…

Computation and Language · Computer Science 2024-12-31 Sanjay Chouhan , Shubha Brata Nath , Aparajita Dutta

Large Language Models (LLMs) play a central role in modern artificial intelligence, yet their development has been primarily focused on English, resulting in limited support for other languages. We present PLLuM (Polish Large Language…

Computation and Language · Computer Science 2025-11-07 Jan Kocoń , Maciej Piasecki , Arkadiusz Janz , Teddy Ferdinan , Łukasz Radliński , Bartłomiej Koptyra , Marcin Oleksy , Stanisław Woźniak , Paweł Walkowiak , Konrad Wojtasik , Julia Moska , Tomasz Naskręt , Bartosz Walkowiak , Mateusz Gniewkowski , Kamil Szyc , Dawid Motyka , Dawid Banach , Jonatan Dalasiński , Ewa Rudnicka , Bartłomiej Alberski , Tomasz Walkowiak , Aleksander Szczęsny , Maciej Markiewicz , Tomasz Bernaś , Hubert Mazur , Kamil Żyta , Mateusz Tykierko , Grzegorz Chodak , Tomasz Kajdanowicz , Przemysław Kazienko , Agnieszka Karlińska , Karolina Seweryn , Anna Kołos , Maciej Chrabąszcz , Katarzyna Lorenc , Aleksandra Krasnodębska , Artur Wilczek , Katarzyna Dziewulska , Paula Betscher , Zofia Cieślińska , Katarzyna Kowol , Daria Mikoś , Maciej Trzciński , Dawid Krutul , Marek Kozłowski , Sławomir Dadas , Rafał Poświata , Michał Perełkiewicz , Małgorzata Grębowiec , Maciej Kazuła , Marcin Białas , Roman Roszko , Danuta Roszko , Jurgita Vaičenonienė , Andrius Utka , Paweł Levchuk , Paweł Kowalski , Irena Prawdzic-Jankowska , Maciej Ogrodniczuk , Monika Borys , Anna Bulińska , Wiktoria Gumienna , Witold Kieraś , Dorota Komosińska , Katarzyna Krasnowska-Kieraś , Łukasz Kobyliński , Martyna Lewandowska , Marek Łaziński , Mikołaj Łątkowski , Dawid Mastalerz , Beata Milewicz , Agnieszka Anna Mykowiecka , Angelika Peljak-Łapińska , Sandra Penno , Zuzanna Przybysz , Michał Rudolf , Piotr Rybak , Karolina Saputa , Aleksandra Tomaszewska , Aleksander Wawer , Marcin Woliński , Joanna Wołoszyn , Alina Wróblewska , Bartosz Żuk , Filip Żarnecki , Konrad Kaczyński , Anna Cichosz , Zuzanna Deckert , Monika Garnys , Izabela Grabarczyk , Wojciech Janowski , Sylwia Karasińska , Aleksandra Kujawiak , Piotr Misztela , Maria Szymańska , Karolina Walkusz , Igor Siek , Jakub Kwiatkowski , Piotr Pęzik

Large Language Models (LLMs) have demonstrated remarkable performance across various natural language tasks, marking significant strides towards general artificial intelligence. While general artificial intelligence is leveraged by…

Computation and Language · Computer Science 2023-10-31 Yizhe Yang , Huashan Sun , Jiawei Li , Runheng Liu , Yinghao Li , Yuhang Liu , Heyan Huang , Yang Gao

Language is essentially a complex, intricate system of human expressions governed by grammatical rules. It poses a significant challenge to develop capable AI algorithms for comprehending and grasping a language. As a major approach,…

Sentiment analysis aims to extract people's emotions and opinion from their comments on the web. It widely used in businesses to detect sentiment in social data, gauge brand reputation, and understand customers. Most of articles in this…

Computation and Language · Computer Science 2022-12-13 Ali Nazarizadeh , Touraj Banirostam , Minoo Sayyadpour

Large language models (LLMs) demonstrate remarkable ability to comprehend, reason, and generate following nature language instructions. However, the development of LLMs has been primarily focused on high-resource languages, such as English,…

Persian music, with its unique tonalities, modal systems (Dastgah), and rhythmic structures, presents significant challenges for music generation models trained primarily on Western music. We address this gap by curating the first…

Recent advancements in text embedding have significantly improved natural language understanding across many languages, yet Persian remains notably underrepresented in large-scale embedding research. In this paper, we present Hakim, a novel…

Computation and Language · Computer Science 2025-10-10 Mehran Sarmadi , Morteza Alikhani , Erfan Zinvandi , Zahra Pourbahman