中文
相关论文

相关论文: Language and Speech Technology for Central Kurdish…

200 篇论文

Discourse parsing is an important task useful for NLU applications such as summarization, machine comprehension, and emotion recognition. The current discourse parsing datasets based on conversations consists of written English dialogues…

计算与语言 · 计算机科学 2025-06-11 Divyaksh Shukla , Ritesh Baviskar , Dwijesh Gohil , Aniket Tiwari , Atul Shree , Ashutosh Modi

Back translation is one of the most widely used methods for improving the performance of neural machine translation systems. Recent research has sought to enhance the effectiveness of this method by increasing the 'diversity' of the…

计算与语言 · 计算机科学 2023-09-01 Laurie Burchell , Alexandra Birch , Kenneth Heafield

Confidently making progress on multilingual modeling requires challenging, trustworthy evaluations. We present TyDi QA---a question answering dataset covering 11 typologically diverse languages with 204K question-answer pairs. The languages…

Large language models (LLMs) have the potential of being useful tools that can automate tasks and assist humans. However, these models are more fluent in English and more aligned with Western cultures, norms, and values. Arabic-specific…

计算与语言 · 计算机科学 2025-03-20 Amr Keleg

Urdu is a challenging language because of, first, its Perso-Arabic script and second, its morphological system having inherent grammatical forms and vocabulary of Arabic, Persian and the native languages of South Asia. This paper describes…

计算与语言 · 计算机科学 2022-04-08 Muhammad Humayoun , Harald Hammarström , Aarne Ranta

Natural language is one of the most fundamental features that distinguish people from other living things and enable people to communicate each other. Language is a tool that enables people to express their feelings and thoughts and to…

计算与语言 · 计算机科学 2019-05-15 Baris Baburoglu , Adem Tekerek , Mehmet Tekerek

We present sentence aligned parallel corpora across 10 Indian Languages - Hindi, Telugu, Tamil, Malayalam, Gujarati, Urdu, Bengali, Oriya, Marathi, Punjabi, and English - many of which are categorized as low resource. The corpora are…

计算与语言 · 计算机科学 2020-07-16 Shashank Siripragada , Jerin Philip , Vinay P. Namboodiri , C V Jawahar

Natural language processing (NLP) has largely focused on modelling standardized languages. More recently, attention has increasingly shifted to local, non-standardized languages and dialects. However, the relevant speaker populations' needs…

计算与语言 · 计算机科学 2024-06-10 Verena Blaschke , Christoph Purschke , Hinrich Schütze , Barbara Plank

Most of the world's languages and dialects are low-resource, and lack support in mainstream machine translation (MT) models. However, many of them have a closely-related high-resource language (HRL) neighbor, and differ in linguistically…

计算与语言 · 计算机科学 2025-10-22 Niyati Bafna , Emily Chang , Nathaniel R. Robinson , David R. Mortensen , Kenton Murray , David Yarowsky , Hale Sirin

Cued Speech (CS) is a visual communication system for the deaf or hearing impaired people. It combines lip movements with hand cues to obtain a complete phonetic repertoire. Current deep learning based methods on automatic CS recognition…

多媒体 · 计算机科学 2021-06-28 Jianrong Wang , Ziyue Tang , Xuewei Li , Mei Yu , Qiang Fang , Li Liu

Zero-shot multi-speaker text-to-speech (ZS-TTS) systems have advanced for English, however, it still lags behind due to insufficient resources. We address this gap for Arabic, a language of more than 450 million native speakers, by first…

计算与语言 · 计算机科学 2024-07-09 Khai Duy Doan , Abdul Waheed , Muhammad Abdul-Mageed

Despite remarkable progress in large language models, Urdu-a language spoken by over 230 million people-remains critically underrepresented in modern NLP systems. Existing multilingual models demonstrate poor performance on Urdu-specific…

计算与语言 · 计算机科学 2026-01-14 Muhammad Taimoor Hassan , Jawad Ahmed , Muhammad Awais

Natural language processing (NLP) and speech technologies have made significant progress in recent years; however, they remain largely focused on standardized language varieties. Dialects, despite their cultural significance and widespread…

计算与语言 · 计算机科学 2026-04-14 Lena S. Oberkircher , Jesujoba O. Alabi , Dietrich Klakow , Jürgen Trouvain

Lipreading has emerged as an increasingly important research area for developing robust speech recognition systems and assistive technologies for the hearing-impaired. However, non-English resources for visual speech recognition remain…

计算机视觉与模式识别 · 计算机科学 2025-10-28 Zahra Taghizadeh , Mohammad Shahverdikondori , Arian Noori , Alireza Dadgarnia

The recent advances in natural language processing have predominantly favored well-resourced English-centric models, resulting in a significant gap with low-resource languages. In this work, we introduce the language model TURNA, which is…

Arabic is a complex language with many varieties and dialects spoken by over 450 millions all around the world. Due to the linguistic diversity and variations, it is challenging to build a robust and generalized ASR system for Arabic. In…

计算与语言 · 计算机科学 2023-10-30 Abdul Waheed , Bashar Talafha , Peter Sullivan , AbdelRahim Elmadany , Muhammad Abdul-Mageed

Multilingual Large Language Models (LLMs) often provide suboptimal performance on low-resource languages like Urdu. This paper introduces UrduLLaMA 1.0, a model derived from the open-source Llama-3.1-8B-Instruct architecture and continually…

计算与语言 · 计算机科学 2025-02-25 Layba Fiaz , Munief Hassan Tahir , Sana Shams , Sarmad Hussain

We investigate structural traces of language contact in the intermediate representations of a monolingual language model. Focusing on Persian (Farsi) as a historically contact-rich language, we probe the representations of a Persian-trained…

计算与语言 · 计算机科学 2026-01-29 Ali Basirat , Danial Namazifard , Navid Baradaran Hemmati

We present the Pashto Common Voice corpus -- the first large-scale, openly licensed speech resource for Pashto, a language with over 60 million native speakers largely absent from open speech technology. Through a community effort spanning…

计算与语言 · 计算机科学 2026-03-31 Hanif Rahman , Shafeeq ur Rehman

This work aims to build a multilingual text-to-speech (TTS) synthesis system for ten lower-resourced Turkic languages: Azerbaijani, Bashkir, Kazakh, Kyrgyz, Sakha, Tatar, Turkish, Turkmen, Uyghur, and Uzbek. We specifically target the…

音频与语音处理 · 电气工程与系统科学 2023-05-26 Rustem Yeshpanov , Saida Mussakhojayeva , Yerbolat Khassanov