中文
相关论文

相关论文: IruMozhi: Automatically classifying diglossia in T…

200 篇论文

As the fourth largest language family in the world, the Dravidian languages have become a research hotspot in natural language processing (NLP). Although the Dravidian languages contain a large number of languages, there are relatively few…

计算与语言 · 计算机科学 2021-12-06 Xiaotian Lin , Nankai Lin , Kanoksak Wattanachote , Shengyi Jiang , Lianxi Wang

We report generation of a MNIST [4] compatible data set [1] for Tamil vowels to enable building a classification DNN or other such ML/AI deep learning [2] models for Tamil OCR/Handwriting applications. We report the capability of the 60,000…

计算机视觉与模式识别 · 计算机科学 2020-06-18 Muthiah Annamalai

While automatic metrics drive progress in Machine Translation (MT) and Text Summarization (TS), existing metrics have been developed and validated almost exclusively for English and other high-resource languages. This narrow focus leaves…

计算与语言 · 计算机科学 2025-10-09 Amir Hossein Yari , Kalmit Kulkarni , Ahmad Raza Khan , Fajri Koto

Language Identification (LI) is crucial for various natural language processing tasks, serving as a foundational step in applications such as sentiment analysis, machine translation, and information retrieval. In multilingual societies like…

计算与语言 · 计算机科学 2025-03-13 Aniket Deroy , Subhankar Maity

In todays digital world automated Machine Translation of one language to another has covered a long way to achieve different kinds of success stories. Whereas Babel Fish supports a good number of foreign languages and only Hindi from Indian…

计算与语言 · 计算机科学 2014-06-17 Siddhartha Ghosh , Sujata Thamke , Kalyani U. R. S

The effectiveness of Large Language Models (LLMs) depends heavily on the availability of high-quality post-training data, particularly instruction-tuning and preference-based examples. Existing open-source datasets, however, often lack…

Neural Machine Translation (NMT) is a predominant machine translation technology nowadays because of its end-to-end trainable flexibility. However, NMT still struggles to translate properly in low-resource settings specifically on distant…

计算与语言 · 计算机科学 2021-09-28 Baban Gain , Dibyanayan Bandyopadhyay , Asif Ekbal

Reduplication and repetition, though similar in form, serve distinct linguistic purposes. Reduplication is a deliberate morphological process used to express grammatical, semantic, or pragmatic nuances, while repetition is often…

计算与语言 · 计算机科学 2024-07-12 Arif Ahmad , Mothika Gayathri Khyathi , Pushpak Bhattacharyya

The rapid proliferation of Large Language Models (LLMs) has created a profound digital divide, effectively excluding indigenous languages of the Global South from the AI revolution. The Tharu language, an Indo-Aryan vernacular spoken by…

计算与语言 · 计算机科学 2026-03-19 Prajwal Panth , Agniva Maiti

Differentiating intrinsic language words from transliterable words is a key step aiding text processing tasks involving different natural languages. We consider the problem of unsupervised separation of transliterable words from native…

计算与语言 · 计算机科学 2018-03-28 Deepak P

In this paper, we survey Text Summarization (TS) datasets in Indian Languages (ILs), which are also low-resource languages (LRLs). We seek to answer one primary question: is the pool of Indian Language Text Summarization (ILTS) dataset…

计算与语言 · 计算机科学 2022-04-28 Shagun Sinha , Girish Nath Jha

India's linguistic landscape is one of the most diverse in the world, comprising over 120 major languages and approximately 1,600 additional languages, with 22 officially recognized as scheduled languages in the Indian Constitution. Despite…

With the growing presence of multilingual users on social media, detecting abusive language in code-mixed text has become increasingly challenging. Code-mixed communication, where users seamlessly switch between English and their native…

计算与语言 · 计算机科学 2025-05-01 Manish Pandey , Nageshwar Prasad Yadav , Mokshada Adduru , Sawan Rai

Natural language processing is a prompt research area across the country. Parsing is one of the very crucial tool in language analysis system which aims to forecast the structural relationship among the words in a given sentence. Many…

计算与语言 · 计算机科学 2014-03-26 K. Sureka , K. G. Srinivasagan , S. Suganthi

We present a novel approach to data preparation for developing multilingual Indic large language model. Our meticulous data acquisition spans open-source and proprietary sources, including Common Crawl, Indic books, news articles, and…

This paper describes how we developed a neural-based dependency parser, namely ThamizhiUDp, which provides a complete pipeline for the dependency parsing of the Tamil language text using Universal Dependency formalism. We have considered…

计算与语言 · 计算机科学 2020-12-29 Kengatharaiyer Sarveswaran , Gihan Dias

Transliteration is a task in the domain of NLP where the output word is a similar-sounding word written using the letters of any foreign language. Today this system has been developed for several language pairs that involve English as…

计算与语言 · 计算机科学 2022-08-24 Yash Raj , Bhavesh Laddagiri

This paper describes the submissions by team HWR to the Dravidian Language Identification (DLI) shared task organized at VarDial 2021 workshop. The DLI training set includes 16,674 YouTube comments written in Roman script containing…

计算与语言 · 计算机科学 2021-03-10 Tommi Jauhiainen , Tharindu Ranasinghe , Marcos Zampieri

Social media platforms often act as breeding grounds for various forms of trolling or malicious content targeting users or communities. One way of trolling users is by creating memes, which in most cases unites an image with a short piece…

多媒体 · 计算机科学 2022-04-28 Mithun Das , Somnath Banerjee , Animesh Mukherjee

We report in this paper, Tamil open-source software community is a vibrant place with software developers, font designers, translators, voice-over artists, and general user testers, who come together for love of their language, and…

计算机与社会 · 计算机科学 2018-01-19 Muthiah Annamalai , T Shrinivasan