中文
相关论文

相关论文: Processing South Asian Languages Written in the La…

200 篇论文

The effectiveness of Large Language Models (LLMs) diminishes for extremely low-resource languages, such as indigenous languages, primarily due to the lack of labeled data. Despite growing interest, the availability of high-quality natural…

计算与语言 · 计算机科学 2026-03-23 Ulin Nuha , Adam Jatowt

Natural Language Processing (NLP) and especially natural language text analysis have seen great advances in recent times. Usage of deep learning in text processing has revolutionized the techniques for text processing and achieved…

信息检索 · 计算机科学 2020-07-07 Ramchandra Joshi , Purvi Goel , Raviraj Joshi

This paper presents a summary of the findings that we obtained based on the shared task on machine translation of Dravidian languages. We stood first in three of the five sub-tasks which were assigned to us for the main shared task. We…

计算与语言 · 计算机科学 2022-04-21 Aditya Vyawahare , Rahul Tangsali , Aditya Mandke , Onkar Litake , Dipali Kadam

The study of historical languages presents unique challenges due to their complex orthographic systems, fragmentary textual evidence, and the absence of standardized digital representations of text in those languages. Tackling these…

Large Language Models (LLMs) are increasingly deployed in high-stakes clinical applications in India. Speakers of Indian languages frequently communicate using romanized text rather than native scripts, yet existing research rarely…

计算与语言 · 计算机科学 2026-04-01 Manurag Khullar , Utkarsh Desai , Poorva Malviya , Aman Dalmia , Zheyuan Ryan Shi

India's vast linguistic diversity presents unique challenges and opportunities for technological advancement, especially in the realm of Natural Language Processing (NLP). While there has been significant progress in NLP applications for…

计算与语言 · 计算机科学 2024-12-25 Rasika Ransing , Mohammed Amaan Dhamaskar , Ayush Rajpurohit , Amey Dhoke , Sanket Dalvi

Recent advances in Deep Learning and Computer Vision have been successfully leveraged to serve marginalized communities in various contexts. One such area is Sign Language - a primary means of communication for the deaf community. However,…

计算机视觉与模式识别 · 计算机科学 2024-01-23 Haz Sameen Shahgir , Khondker Salman Sayeed , Md Toki Tahmid , Tanjeem Azwad Zaman , Md. Zarif Ul Alam

We introduce Jambu, a cognate database of South Asian languages which unifies dozens of previous sources in a structured and accessible format. The database includes 287k lemmata from 602 lects, grouped together in 23k sets of cognates. We…

计算与语言 · 计算机科学 2023-06-06 Aryaman Arora , Adam Farris , Samopriya Basu , Suresh Kolichala

This paper introduces \textit{Bangla Key2Text}, a large-scale dataset of $2.6$ million Bangla keyword--text pairs designed for keyword-driven text generation in a low-resource language. The dataset is constructed using a BERT-based keyword…

计算与语言 · 计算机科学 2026-04-22 Tonmoy Talukder , G M Shahariar

We present CLASSLA-Stanza, a pipeline for automatic linguistic annotation of the South Slavic languages, which is based on the Stanza natural language processing pipeline. We describe the main improvements in CLASSLA-Stanza with respect to…

计算与语言 · 计算机科学 2023-08-14 Luka Terčon , Nikola Ljubešić

Cross-lingual summarization (CLS) is the task to produce a summary in one particular language for a source document in a different language. We introduce WikiMulti - a new dataset for cross-lingual summarization based on Wikipedia articles…

计算与语言 · 计算机科学 2022-04-26 Pavel Tikhonov , Valentin Malykh

People commonly communicate in English, Arabic, and Bengali spoken languages through various mediums. However, deaf and hard-of-hearing individuals primarily use body language and sign language to express their needs and achieve…

计算机视觉与模式识别 · 计算机科学 2024-08-21 Md Hadiuzzaman , Mohammed Sowket Ali , Tamanna Sultana , Abdur Raj Shafi , Abu Saleh Musa Miah , Jungpil Shin

We present sentence aligned parallel corpora across 10 Indian Languages - Hindi, Telugu, Tamil, Malayalam, Gujarati, Urdu, Bengali, Oriya, Marathi, Punjabi, and English - many of which are categorized as low resource. The corpora are…

计算与语言 · 计算机科学 2020-07-16 Shashank Siripragada , Jerin Philip , Vinay P. Namboodiri , C V Jawahar

The national languages of Senegal, like those of West Africa country in general, are written with two alphabets : the Latin alphabet that draws its strength from official decreesm and the completed Arabic script (Ajami), widespread and well…

计算与语言 · 计算机科学 2020-05-07 El hadji M. Fall , El hadji M. Nguer , Bao Diop Sokhna , Mouhamadou Khoule , Mathieu Mangeot , Mame T. Cisse

Automated text simplification aims to produce simple versions of complex texts. This task is especially useful in the medical domain, where the latest medical findings are typically communicated via complex and technical articles. This…

Developing culturally grounded multilingual AI systems remains challenging, particularly for low-resource languages. While synthetic data offers promise, its effectiveness in multilingual and multicultural contexts is underexplored. We…

We present V\=arta, a large-scale multilingual dataset for headline generation in Indic languages. This dataset includes 41.8 million news articles in 14 different Indic languages (and English), which come from a variety of high-quality…

计算与语言 · 计算机科学 2023-05-11 Rahul Aralikatte , Ziling Cheng , Sumanth Doddapaneni , Jackie Chi Kit Cheung

Scene text recognition is essential in many applications, including automated translation, information retrieval, driving assistance, and enhancing accessibility for individuals with visual impairments. Much research has been done to…

计算机视觉与模式识别 · 计算机科学 2024-05-21 Fadila Wendigoundi Douamba , Jianjun Song , Ling Fu , Yuliang Liu , Xiang Bai

MIRACL (Multilingual Information Retrieval Across a Continuum of Languages) is a multilingual dataset we have built for the WSDM 2023 Cup challenge that focuses on ad hoc retrieval across 18 different languages, which collectively encompass…

Arabic is recognised as the 4th most used language of the Internet. Arabic has three main varieties: (1) classical Arabic (CA), (2) Modern Standard Arabic (MSA), (3) Arabic Dialect (AD). MSA and AD could be written either in Arabic or in…

计算与语言 · 计算机科学 2019-03-08 Imane Guellil , Houda Saâdane , Faical Azouaou , Billel Gueni , Damien Nouvel
‹ 上一页 1 8 9 10 下一页 ›