中文
相关论文

相关论文: Automatic Identification of Closely-related Indian…

200 篇论文

In this paper we present a system based on SVM ensembles trained on characters and words to discriminate between five similar languages of the Indo-Aryan family: Hindi, Braj Bhasha, Awadhi, Bhojpuri, and Magahi. We investigate the…

计算与语言 · 计算机科学 2018-07-10 Alina Maria Ciobanu , Marcos Zampieri , Shervin Malmasi , Santanu Pal , Liviu P. Dinu

Language identification has become a prerequisite for all kinds of automated text processing systems. In this paper, we present a rule-based language identifier tool for two closely related Indo-Aryan languages: Hindi and Magahi. This…

计算与语言 · 计算机科学 2018-04-17 Priya Rani , Atul Kr. Ojha , Girish Nath Jha

In this paper we discuss an in-progress work on the development of a speech corpus for four low-resource Indo-Aryan languages -- Awadhi, Bhojpuri, Braj and Magahi using the field methods of linguistic data collection. The total size of the…

Language Identification is a very important part of several text processing pipelines. Extensive research has been done in this field. This paper proposes a procedure for automatic language identification of poems for poem analysis task,…

计算与语言 · 计算机科学 2021-01-01 Priyankit Acharya , Aditya Ku. Pathak , Rakesh Ch. Balabantaray , Anil Ku. Singh

Corpus preparation for low-resource languages and for development of human language technology to analyze or computationally process them is a laborious task, primarily due to the unavailability of expert linguists who are native speakers…

计算与语言 · 计算机科学 2021-08-18 Rajesh Kumar Mundotiya , Manish Kumar Singh , Rahul Kapur , Swasti Mishra , Anil Kumar Singh

We present sentence aligned parallel corpora across 10 Indian Languages - Hindi, Telugu, Tamil, Malayalam, Gujarati, Urdu, Bengali, Oriya, Marathi, Punjabi, and English - many of which are categorized as low resource. The corpora are…

计算与语言 · 计算机科学 2020-07-16 Shashank Siripragada , Jerin Philip , Vinay P. Namboodiri , C V Jawahar

This paper focuses on developing translation models and related applications for 36 Indian languages, including Assamese, Awadhi, Bengali, Bhojpuri, Braj, Bodo, Dogri, English, Konkani, Gondi, Gujarati, Hindi, Hinglish, Ho, Kannada, Kangri,…

计算与语言 · 计算机科学 2025-01-03 Vandan Mujadia , Dipti Misra Sharma

Automatic Cognate Detection (ACD) is a challenging task which has been utilized to help NLP applications like Machine Translation, Information Retrieval and Computational Phylogenetics. Unidentified cognate pairs can pose a challenge to…

计算与语言 · 计算机科学 2022-01-03 Diptesh Kanojia , Kevin Patel , Pushpak Bhattacharyya , Malhar Kulkarni , Gholamreza Haffari

In this work, we present an extensive study of statistical machine translation involving languages of the Indian subcontinent. These languages are related by genetic and contact relationships. We describe the similarities between Indic…

计算与语言 · 计算机科学 2020-03-20 Anoop Kunchukuttan , Pushpak Bhattacharyya

Spoken language Identification (LID) systems are needed to identify the language(s) present in a given audio sample, and typically could be the first step in many speech processing related tasks such as automatic speech recognition (ASR).…

计算与语言 · 计算机科学 2020-10-15 Pradeep Rangan , Sundeep Teki , Hemant Misra

Most existing approaches for unsupervised bilingual lexicon induction (BLI) depend on good quality static or contextual embeddings requiring large monolingual corpora for both languages. However, unsupervised BLI is most likely to be useful…

计算与语言 · 计算机科学 2024-03-26 Niyati Bafna , Cristina España-Bonet , Josef van Genabith , Benoît Sagot , Rachel Bawden

Automatic speech recognition (ASR) performance has improved drastically in recent years, mainly enabled by self-supervised learning (SSL) based acoustic models such as wav2vec2 and large-scale multi-lingual training like Whisper. A huge…

Large language models and multilingual machine translation (MT) systems increasingly drive access to information, yet many languages of the tribal communities remain effectively invisible in these technologies. This invisibility exacerbates…

计算与语言 · 计算机科学 2025-12-05 Pooja Singh , Sandeep Kumar

We explore the impact of leveraging the relatedness of languages that belong to the same family in NLP models using multilingual fine-tuning. We hypothesize and validate that multilingual fine-tuning of pre-trained language models can yield…

We create publicly available language identification (LID) datasets and models in all 22 Indian languages listed in the Indian constitution in both native-script and romanized text. First, we create Bhasha-Abhijnaanam, a language…

计算与语言 · 计算机科学 2023-10-27 Yash Madhani , Mitesh M. Khapra , Anoop Kunchukuttan

Language identification is used as the first step in many data collection and crawling efforts because it allows us to sort online text into language-specific buckets. However, many modern languages, such as Konkani, Kashmiri, Punjabi etc.,…

计算与语言 · 计算机科学 2024-06-27 Milind Agarwal , Joshua Otten , Antonios Anastasopoulos

Magahi is an Indo-Aryan Language, spoken mainly in the Eastern parts of India. Despite having a significant number of speakers, there has been virtually no language resource (LR) or language technology (LT) developed for the language,…

计算与语言 · 计算机科学 2021-12-01 Ritesh Kumar

In this paper, we conduct one of the very first studies for cross-corpora performance evaluation in the spoken language identification (LID) problem. Cross-corpora evaluation was not explored much in LID research, especially for the Indian…

音频与语音处理 · 电气工程与系统科学 2021-05-13 Spandan Dey , Goutam Saha , Md Sahidullah

The language identification task is a crucial fundamental step in NLP. Often it serves as a pre-processing step for widely used NLP applications such as multilingual machine translation, information retrieval, question and answering, and…

计算与语言 · 计算机科学 2026-01-08 Yash Ingle , Pruthwik Mishra

Automatic spoken language identification (LID) is a very important research field in the era of multilingual voice-command-based human-computer interaction (HCI). A front-end LID module helps to improve the performance of many speech-based…

计算与语言 · 计算机科学 2022-12-08 Spandan Dey , Md Sahidullah , Goutam Saha
‹ 上一页 1 2 3 10 下一页 ›