中文
相关论文

相关论文: Automatic Language Identification System for Hindi…

200 篇论文

Evaluating instruction-tuned Large Language Models (LLMs) in Hindi is challenging due to a lack of high-quality benchmarks, as direct translation of English datasets fails to capture crucial linguistic and cultural nuances. To address this,…

India is a multilingual society with 1369 rationalized languages and dialects being spoken across the country (INDIA, 2011). Of these, the 22 scheduled languages have a staggering total of 1.17 billion speakers and 121 languages have more…

We explore the impact of leveraging the relatedness of languages that belong to the same family in NLP models using multilingual fine-tuning. We hypothesize and validate that multilingual fine-tuning of pre-trained language models can yield…

Language Identification (LID) is a challenging task, especially when the input texts are short and noisy such as posts and statuses on social media or chat logs on gaming forums. The task has been tackled by either designing a feature set…

计算与语言 · 计算机科学 2019-10-16 Duy Tin Vo , Richard Khoury

This paper provides an overall introduction of our Automatic Speech Recognition (ASR) systems for Southeast Asian languages. As not much existing work has been carried out on such regional languages, a few difficulties should be addressed…

计算与语言 · 计算机科学 2022-10-10 Lei Wang , Rong Tong , Cheung Chi Leung , Sunil Sivadas , Chongjia Ni , Bin Ma

The aim of this paper is to develop a flexible framework capable of automatically recognizing phonetic units present in a speech utterance of any language spoken in any mode. In this study, we considered two modes of speech: conversation,…

音频与语音处理 · 电气工程与系统科学 2019-08-27 Kumud Tripathi , M. Kiran Reddy , K. Sreenivasa Rao

Mining parallel document pairs for document-level machine translation (MT) remains challenging due to the limitations of existing Cross-Lingual Document Alignment (CLDA) techniques. Existing methods often rely on metadata such as URLs,…

计算与语言 · 计算机科学 2025-11-11 Sanjay Suryanarayanan , Haiyue Song , Mohammed Safi Ur Rahman Khan , Anoop Kunchukuttan , Raj Dabre

Language identification (LID) is a fundamental step in many natural language processing pipelines. However, current LID systems are far from perfect, particularly on lower-resource languages. We present a LID model which achieves a…

计算与语言 · 计算机科学 2023-08-31 Laurie Burchell , Alexandra Birch , Nikolay Bogoychev , Kenneth Heafield

Tamil, a Dravidian language of South Asia, is a highly diglossic language with two very different registers in everyday use: Literary Tamil (preferred in writing and formal communication) and Spoken Tamil (confined to speech and informal…

计算与语言 · 计算机科学 2023-11-15 Kabilan Prasanna , Aryaman Arora

India's linguistic landscape is one of the most diverse in the world, comprising over 120 major languages and approximately 1,600 additional languages, with 22 officially recognized as scheduled languages in the Indian Constitution. Despite…

Multilingual ASR technology simplifies model training and deployment, but its accuracy is known to depend on the availability of language information at runtime. Since language identity is seldom known beforehand in real-world scenarios, it…

Automatic assessment of reading fluency using automatic speech recognition (ASR) holds great potential for early detection of reading difficulties and subsequent timely intervention. Precise assessment tools are required, especially for…

计算与语言 · 计算机科学 2024-07-24 Bo Molenaar , Cristian Tejedor-Garcia , Helmer Strik , Catia Cucchiarini

This paper describes the system submitted to Dravidian-Codemix-HASOC2021: Hate Speech and Offensive Language Identification in Dravidian Languages (Tamil-English and Malayalam-English). This task aims to identify offensive content in…

计算与语言 · 计算机科学 2021-12-08 Sean Benhur , Kanchana Sivanraju

Ragas form the foundation for Indian Classical Music. The task of Raga Recognition has gained traction in the Music Information Retrieval community in the recent past, which can be attributed to the nuances of Indian Classical Music that…

声音 · 计算机科学 2022-12-13 Devayani Hebbar , Vandana Jagtap

Most legal text in the Indian judiciary is written in complex English due to historical reasons. However, only a small fraction of the Indian population is comfortable in reading English. Hence legal text needs to be made available in…

计算与语言 · 计算机科学 2024-11-08 Sayan Mahapatra , Debtanu Datta , Shubham Soni , Adrijit Goswami , Saptarshi Ghosh

In this work, we describe a system that detects paraphrases in Indian Languages as part of our participation in the shared Task on detecting paraphrases in Indian Languages (DPIL) organized by Forum for Information Retrieval Evaluation…

计算与语言 · 计算机科学 2016-12-28 Kamal Sarkar

Multilingual large language models (LLMs) are increasingly deployed in linguistically diverse regions like India, yet most interpretability tools remain tailored to English. Prior work reveals that LLMs often operate in English centric…

计算与语言 · 计算机科学 2026-02-19 Mihir Panchal , Deeksha Varshney , Mamta , Asif Ekbal

It is a well-known fact that current AI-based language technology -- language models, machine translation systems, multilingual dictionaries and corpora -- focuses on the world's 2-3% most widely spoken languages. Recent research efforts…

计算与语言 · 计算机科学 2023-07-26 Gábor Bella , Paula Helm , Gertraud Koch , Fausto Giunchiglia

Spoken Language Identification (LID) is an important sub-task of Automatic Speech Recognition(ASR) that is used to classify the language(s) in an audio segment. Automatic LID plays an useful role in multilingual countries. In various…

音频与语音处理 · 电气工程与系统科学 2024-09-02 Parth Shastri , Chirag Patil , Poorval Wanere , Shrinivas Mahajan , Abhishek Bhatt , Hardik Sailor

Sign language, which contains hand movements, facial expressions and bodily gestures, is a significant medium for communicating with hard-of-hearing people. A well-trained sign language community communicates easily, but those who don't…

计算机视觉与模式识别 · 计算机科学 2025-08-14 Ajeet Kumar Yadav , Nishant Kumar , Rathna G N