中文
相关论文

相关论文: Discriminating between Indo-Aryan Languages Using …

200 篇论文

Cognates are variants of the same lexical form across different languages; for example 'fonema' in Spanish and 'phoneme' in English are cognates, both of which mean 'a unit of sound'. The task of automatic detection of cognates among any…

计算与语言 · 计算机科学 2021-12-17 Diptesh Kanojia , Raj Dabre , Shubham Dewangan , Pushpak Bhattacharyya , Gholamreza Haffari , Malhar Kulkarni

Quite often, words from one language are adopted within a different language without translation; these words appear in transliterated form in text written in the latter language. This phenomenon is particularly widespread within Indian…

计算与语言 · 计算机科学 2020-05-07 Sridhama Prakhya , Deepak P

Being less resource languages, Indian-Indian and English-Indian language MT system developments faces the difficulty to translate various lexical phenomena. In this paper, we present our work on a comparative study of 440 phrase-based…

计算与语言 · 计算机科学 2017-10-09 Sreelekha S , Pushpak Bhattacharyya

This paper describes the system submitted to Dravidian-Codemix-HASOC2021: Hate Speech and Offensive Language Identification in Dravidian Languages (Tamil-English and Malayalam-English). This task aims to identify offensive content in…

计算与语言 · 计算机科学 2021-12-08 Sean Benhur , Kanchana Sivanraju

The evolution of languages closely resembles the evolution of haploid organisms. This similarity has been recently exploited \cite{GA,GJ} to construct language trees. The key point is the definition of a distance among all pairs of…

物理与社会 · 物理学 2009-11-13 Maurizio Serva , Filippo Petroni

Translating technical terms into lexically similar, low-resource Indian languages remains a challenge due to limited parallel data and the complexity of linguistic structures. We propose a novel use-case of Sanskrit-based segments for…

计算与语言 · 计算机科学 2026-03-26 Karthika N J , Krishnakant Bhatt , Ganesh Ramakrishnan , Preethi Jyothi

People with vocal and hearing disabilities use sign language to express themselves using visual gestures and signs. Although sign language is a solution for communication difficulties faced by deaf people, there are still problems as most…

计算机视觉与模式识别 · 计算机科学 2023-05-01 Mallikharjuna Rao K , Harleen Kaur , Sanjam Kaur Bedi , M A Lekhana

The Ramayana is among the most influential literary traditions of South and Southeast Asia, transmitted across numerous linguistic and cultural contexts over two millennia. Despite extensive scholarship on regional Ramayana traditions,…

计算与语言 · 计算机科学 2026-04-16 Sumesh VP

Recent advances in Unsupervised Neural Machine Translation (UNMT) have minimized the gap between supervised and unsupervised machine translation performance for closely related language pairs. However, the situation is very different for…

计算与语言 · 计算机科学 2021-06-10 Tamali Banerjee , Rudra Murthy , Pushpak Bhattacharyya

This paper presents a summary of the findings that we obtained based on the shared task on machine translation of Dravidian languages. We stood first in three of the five sub-tasks which were assigned to us for the main shared task. We…

计算与语言 · 计算机科学 2022-04-21 Aditya Vyawahare , Rahul Tangsali , Aditya Mandke , Onkar Litake , Dipali Kadam

In a multilingual country like India, multilingual Automatic Speech Recognition (ASR) systems have much scope. Multilingual ASR systems exhibit many advantages like scalability, maintainability, and improved performance over the monolingual…

音频与语音处理 · 电气工程与系统科学 2022-11-01 Arunkumar A , Mudit Batra , Umesh S

This study explores the use of self-supervised learning (SSL) models for tone recognition in three low-resource languages from North Eastern India: Angami, Ao, and Mizo. We evaluate four Wav2vec2.0 base models that were pre-trained on both…

音频与语音处理 · 电气工程与系统科学 2025-06-05 Parismita Gogoi , Sishir Kalita , Wendy Lalhminghlui , Viyazonuo Terhiija , Moakala Tzudir , Priyankoo Sarmah , S. R. M. Prasanna

Language identification of social media text still remains a challenging task due to properties like code-mixing and inconsistent phonetic transliterations. In this paper, we present a supervised learning approach for language…

计算与语言 · 计算机科学 2018-06-28 Soumil Mandal , Sourya Dipta Das , Dipankar Das

The rapid advancement of large language models (LLMs) necessitates evaluation frameworks that reflect real-world academic rigor and multilingual complexity. This paper introduces IndicEval, a scalable benchmarking platform designed to…

计算与语言 · 计算机科学 2026-02-19 Saurabh Bharti , Gaurav Azad , Abhinaw Jagtap , Nachiket Tapas

The Latin script is often used to informally write languages with non-Latin native scripts. In many cases (e.g., most languages in India), the lack of conventional spelling in the Latin script results in high spelling variability. Such…

计算与语言 · 计算机科学 2025-11-19 Adrian Benton , Alexander Gutkin , Christo Kirov , Brian Roark

This paper presents our modeling and architecture approaches for building a highly accurate low-latency language identification system to support multilingual spoken queries for voice assistants. A common approach to solve multilingual…

音频与语音处理 · 电气工程与系统科学 2020-06-02 Chander Chandak , Zeynab Raeesy , Ariya Rastrow , Yuzong Liu , Xiangyang Huang , Siyu Wang , Dong Kwon Joo , Roland Maas

Large Language Models (LLMs) perform well on unseen tasks in English, but their abilities in non English languages are less explored due to limited benchmarks and training data. To bridge this gap, we introduce the Indic QA Benchmark, a…

In a multilingual or sociolingual configuration Intra-sentential Code Switching (ICS) or Code Mixing (CM) is frequently observed nowadays. In the world, most of the people know more than one language. CM usage is especially apparent in…

计算与语言 · 计算机科学 2020-10-12 Sunil Gundapu , Radhika Mamidi

Voice activity detection (VAD), used as the front end of speech enhancement, speech and speaker recognition algorithms, determines the overall accuracy and efficiency of the algorithms. Therefore, a VAD with low complexity and high accuracy…

声音 · 计算机科学 2019-02-06 Jayanta Dey , Md Sanzid Bin Hossain , Mohammad Ariful Haque

The selection of features for text classification is a fundamental task in text mining and information retrieval. Despite being the sixth most widely spoken language in the world, Bangla has received little attention due to the scarcity of…

信息检索 · 计算机科学 2023-08-29 Md. Rafi-Ur-Rashid , Sami Azam , Mirjam Jonkman