中文
相关论文

相关论文: Language Resources and Technologies for Non-Schedu…

200 篇论文

India is country of several hundred different languages. Though twenty two languages have only been devised as scheduled to the Eighth Schedule of Indian Constitution in 2007. But as there is yet no proposed compact display architecture to…

其他计算机科学 · 计算机科学 2012-08-06 Partha Pratim Ray

India has a rich linguistic landscape with languages from 4 major language families spoken by over a billion people. 22 of these languages are listed in the Constitution of India (referred to as scheduled languages) are the focus of this…

India's linguistic landscape, spanning 22 scheduled languages and hundreds of marginalized dialects, has driven rapid growth in NLP datasets, benchmarks, and pretrained models. However, no dedicated survey consolidates resources developed…

计算与语言 · 计算机科学 2026-04-21 Raghvendra Kumar , Devankar Raj , Sriparna Saha

India is home to a multitude of languages of which 22 languages are recognised by the Indian Constitution as official. Building speech based applications for the Indian population is a difficult problem owing to limited data and the number…

According to UNESCO, there are nearly 7,000 languages spoken worldwide, of which around 3,000 languages are in danger of disappearing before the end of the century. With roughly 230 languages having already become extinct between the years…

人机交互 · 计算机科学 2023-04-20 Dinesh Kumar Nanduri , Elizabeth M. Bonsignore

Recent methods in speech and language technology pretrain very LARGE models which are fine-tuned for specific tasks. However, the benefits of such LARGE models are often limited to a few resource rich languages of the world. In this work,…

We present Vakyansh, an end to end toolkit for Speech Recognition in Indic languages. India is home to almost 121 languages and around 125 crore speakers. Yet most of the languages are low resource in terms of data and pretrained models.…

India has 1369 languages of which 22 are official. About 13 different scripts are used to represent these languages. A Common Label Set (CLS) was developed based on phonetics to address the issue of large vocabulary of units required in the…

计算与语言 · 计算机科学 2025-02-24 Utkarsh P

India's linguistic diversity presents both opportunities and challenges for fintech platforms. While the country has 31 major languages and over 100 minor ones, only 10\% of the population understands English, creating barriers to financial…

Language technologies contribute to promoting multilingualism and linguistic diversity around the world. However, only a very small number of the over 7000 languages of the world are represented in the rapidly evolving language technologies…

计算与语言 · 计算机科学 2021-01-28 Pratik Joshi , Sebastin Santy , Amar Budhiraja , Kalika Bali , Monojit Choudhury

We present sentence aligned parallel corpora across 10 Indian Languages - Hindi, Telugu, Tamil, Malayalam, Gujarati, Urdu, Bengali, Oriya, Marathi, Punjabi, and English - many of which are categorized as low resource. The corpora are…

计算与语言 · 计算机科学 2020-07-16 Shashank Siripragada , Jerin Philip , Vinay P. Namboodiri , C V Jawahar

Natural Language processing (NLP) represents the task of automatic handling of natural human language by machines.There is large spectrum of possible applications of NLP which help in automating tasks like translating text from one language…

计算与语言 · 计算机科学 2021-02-02 Nikita P. Desai , Prof. , Vipul K. Dabhi

Magahi is an Indo-Aryan Language, spoken mainly in the Eastern parts of India. Despite having a significant number of speakers, there has been virtually no language resource (LR) or language technology (LT) developed for the language,…

计算与语言 · 计算机科学 2021-12-01 Ritesh Kumar

Languages are classified as low-resource when they lack the quantity of data necessary for training statistical and machine learning tools and models. Causes of resource scarcity vary but can include poor access to technology for developing…

计算与语言 · 计算机科学 2022-04-13 Zoey Liu , Crystal Richardson , Richard Hatcher , Emily Prud'hommeaux

The primary obstacle to developing technologies for low-resource languages is the lack of usable data. In this paper, we report the adoption and deployment of 4 technology-driven methods of data collection for Gondi, a low-resource…

This report evaluates the performance of text-in text-out Large Language Models (LLMs) to understand and generate Indic languages. This evaluation is used to identify and prioritize Indic languages suited for inclusion in safety benchmarks.…

计算与语言 · 计算机科学 2025-01-24 Aatman Vaidya , Tarunima Prabhakar , Denny George , Swair Shah

Despite the considerable advancements in English LLMs, the progress in building comparable models for other languages has been hindered due to the scarcity of tailored resources. Our work aims to bridge this divide by introducing an…

India is a multilingual society with 1369 rationalized languages and dialects being spoken across the country (INDIA, 2011). Of these, the 22 scheduled languages have a staggering total of 1.17 billion speakers and 121 languages have more…

In this paper, we examine and analyze the challenges associated with developing and introducing language technologies to low-resource language communities. While doing so, we bring to light the successes and failures of past work in this…

Existing Indic ASR benchmarks often use scripted, clean speech and leaderboard driven evaluation that encourages dataset specific overfitting. In addition, strict single reference WER penalizes natural spelling variation in Indian…

‹ 上一页 1 2 3 10 下一页 ›