中文
相关论文

相关论文: Discriminating between Indo-Aryan Languages Using …

200 篇论文

In this paper, we discuss an attempt to develop an automatic language identification system for 5 closely-related Indo-Aryan languages of India, Awadhi, Bhojpuri, Braj, Hindi and Magahi. We have compiled a comparable corpora of varying…

计算与语言 · 计算机科学 2018-03-28 Ritesh Kumar , Bornini Lahiri , Deepak Alok , Atul Kr. Ojha , Mayank Jain , Abdul Basit , Yogesh Dawer

Language identification has become a prerequisite for all kinds of automated text processing systems. In this paper, we present a rule-based language identifier tool for two closely related Indo-Aryan languages: Hindi and Magahi. This…

计算与语言 · 计算机科学 2018-04-17 Priya Rani , Atul Kr. Ojha , Girish Nath Jha

Language Identification is a very important part of several text processing pipelines. Extensive research has been done in this field. This paper proposes a procedure for automatic language identification of poems for poem analysis task,…

计算与语言 · 计算机科学 2021-01-01 Priyankit Acharya , Aditya Ku. Pathak , Rakesh Ch. Balabantaray , Anil Ku. Singh

In this paper we discuss an in-progress work on the development of a speech corpus for four low-resource Indo-Aryan languages -- Awadhi, Bhojpuri, Braj and Magahi using the field methods of linguistic data collection. The total size of the…

This paper focuses on developing translation models and related applications for 36 Indian languages, including Assamese, Awadhi, Bengali, Bhojpuri, Braj, Bodo, Dogri, English, Konkani, Gondi, Gujarati, Hindi, Hinglish, Ho, Kannada, Kangri,…

计算与语言 · 计算机科学 2025-01-03 Vandan Mujadia , Dipti Misra Sharma

In this paper we present the GDI_classification entry to the second German Dialect Identification (GDI) shared task organized within the scope of the VarDial Evaluation Campaign 2018. We present a system based on SVM classifier ensembles…

计算与语言 · 计算机科学 2018-07-24 Alina Maria Ciobanu , Shervin Malmasi , Liviu P. Dinu

In this paper we present ensemble-based systems for dialect and language variety identification using the datasets made available by the organizers of the VarDial Evaluation Campaign 2018. We present a system developed to discriminate…

计算与语言 · 计算机科学 2018-08-15 Liviu P. Dinu , Alina Maria Ciobanu , Marcos Zampieri , Shervin Malmasi

We explore the impact of leveraging the relatedness of languages that belong to the same family in NLP models using multilingual fine-tuning. We hypothesize and validate that multilingual fine-tuning of pre-trained language models can yield…

Improving ASR systems is necessary to make new LLM-based use-cases accessible to people across the globe. In this paper, we focus on Indian languages, and make the case that diverse benchmarks are required to evaluate and improve ASR…

计算与语言 · 计算机科学 2023-08-03 Kaushal Santosh Bhogale , Sai Sundaresan , Abhigyan Raman , Tahir Javed , Mitesh M. Khapra , Pratyush Kumar

This paper presents an ensemble system combining the output of multiple SVM classifiers to native language identification (NLI). The system was submitted to the NLI Shared Task 2017 fusion track which featured students essays and spoken…

计算与语言 · 计算机科学 2017-07-25 Marcos Zampieri , Alina Maria Ciobanu , Liviu P. Dinu

Spoken language Identification (LID) systems are needed to identify the language(s) present in a given audio sample, and typically could be the first step in many speech processing related tasks such as automatic speech recognition (ASR).…

计算与语言 · 计算机科学 2020-10-15 Pradeep Rangan , Sundeep Teki , Hemant Misra

Corpus preparation for low-resource languages and for development of human language technology to analyze or computationally process them is a laborious task, primarily due to the unavailability of expert linguists who are native speakers…

计算与语言 · 计算机科学 2021-08-18 Rajesh Kumar Mundotiya , Manish Kumar Singh , Rahul Kapur , Swasti Mishra , Anil Kumar Singh

The recent surge of complex attention-based deep learning architectures has led to extraordinary results in various downstream NLP tasks in the English language. However, such research for resource-constrained and morphologically rich…

计算与语言 · 计算机科学 2021-02-23 Atharva Kulkarni , Amey Hengle , Rutuja Udyawar

Automatic speech recognition (ASR) performance has improved drastically in recent years, mainly enabled by self-supervised learning (SSL) based acoustic models such as wav2vec2 and large-scale multi-lingual training like Whisper. A huge…

Large language models and multilingual machine translation (MT) systems increasingly drive access to information, yet many languages of the tribal communities remain effectively invisible in these technologies. This invisibility exacerbates…

计算与语言 · 计算机科学 2025-12-05 Pooja Singh , Sandeep Kumar

Social media platforms serve as accessible outlets for individuals to express their thoughts and experiences, resulting in an influx of user-generated data spanning all age groups. While these platforms enable free expression, they also…

计算与语言 · 计算机科学 2023-12-12 Nikhil Narayan , Mrutyunjay Biswal , Pramod Goyal , Abhranta Panigrahi

We present Vakyansh, an end to end toolkit for Speech Recognition in Indic languages. India is home to almost 121 languages and around 125 crore speakers. Yet most of the languages are low resource in terms of data and pretrained models.…

This paper describes the systems developed by SPRING Lab, Indian Institute of Technology Madras, for the ASRU MADASR 2.0 challenge. The systems developed focuses on adapting ASR systems to improve in predicting the language and dialect of…

计算与语言 · 计算机科学 2025-11-20 Arjun Gangwar , Kaousheik Jayakumar , S. Umesh

Speaker Verification (SV) is a task to verify the claimed identity of the claimant using his/her voice sample. Though there exists an ample amount of research in SV technologies, the development concerning a multilingual conversation is…

音频与语音处理 · 电气工程与系统科学 2023-02-28 Jagabandhu Mishra , Mrinmoy Bhattacharjee , S. R. Mahadeva Prasanna

Training a conventional automatic speech recognition (ASR) system to support multiple languages is challenging because the sub-word unit, lexicon and word inventories are typically language specific. In contrast, sequence-to-sequence models…

音频与语音处理 · 电气工程与系统科学 2018-02-16 Shubham Toshniwal , Tara N. Sainath , Ron J. Weiss , Bo Li , Pedro Moreno , Eugene Weinstein , Kanishka Rao
‹ 上一页 1 2 3 10 下一页 ›