中文
相关论文

相关论文: ILID: Native Script Language Identification for In…

200 篇论文

Native Language Identification (NLI) intends to classify an author's native language based on their writing in another language. Historically, the task has heavily relied on time-consuming linguistic feature engineering, and…

计算与语言 · 计算机科学 2023-09-14 Sergey Kramp , Giovanni Cassani , Chris Emmery

India is a diverse society with unique challenges in developing AI systems, including linguistic diversity, oral traditions, data accessibility, and scalability. Existing foundation models are primarily trained on English, limiting their…

This paper describes the submissions by team HWR to the Dravidian Language Identification (DLI) shared task organized at VarDial 2021 workshop. The DLI training set includes 16,674 YouTube comments written in Roman script containing…

计算与语言 · 计算机科学 2021-03-10 Tommi Jauhiainen , Tharindu Ranasinghe , Marcos Zampieri

In a multilingual country like India where 12 different official scripts are in use, automatic identification of handwritten script facilitates many important applications such as automatic transcription of multilingual documents, searching…

计算机视觉与模式识别 · 计算机科学 2020-09-17 Pawan Kumar Singh , Iman Chatterjee , Ram Sarkar , Mita Nasipuri

This paper presents a detailed system description of our entry for the CHiPSAL 2025 shared task, focusing on language detection, hate speech identification, and target detection in Devanagari script languages. We experimented with a…

ChatGPT has recently emerged as a powerful NLP tool that can carry out a variety of tasks. However, the range of languages ChatGPT can handle remains largely a mystery. To uncover which languages ChatGPT `knows', we investigate its language…

计算与语言 · 计算机科学 2024-04-10 Wei-Rui Chen , Ife Adebara , Khai Duy Doan , Qisheng Liao , Muhammad Abdul-Mageed

Large language models (LLMs) are increasingly applied in multilingual contexts, yet their capacity for consistent, logically grounded alignment across languages remains underexplored. We present a controlled evaluation framework for…

计算与语言 · 计算机科学 2025-08-21 Samir Abdaljalil , Erchin Serpedin , Khalid Qaraqe , Hasan Kurban

India's rich cultural and linguistic diversity poses various challenges in the domain of Natural Language Processing (NLP), particularly in Named Entity Recognition (NER). NER is a NLP task that aims to identify and classify tokens into…

计算与语言 · 计算机科学 2025-02-07 Mohammed Amaan Dhamaskar , Rasika Ransing

Instruction-following benchmarks remain predominantly English-centric, leaving a critical evaluation gap for the hundreds of millions of Indic language speakers. We introduce IndicIFEval, a benchmark evaluating constrained generation of…

计算与语言 · 计算机科学 2026-02-26 Thanmay Jayakumar , Mohammed Safi Ur Rahman Khan , Raj Dabre , Ratish Puduppully , Anoop Kunchukuttan

Given the large number of Hindi speakers worldwide, there is a pressing need for robust and efficient information retrieval systems for Hindi. Despite ongoing research, there is a lack of comprehensive benchmark for evaluating retrieval…

信息检索 · 计算机科学 2024-08-20 Arkadeep Acharya , Rudra Murthy , Vishwajeet Kumar , Jaydeep Sen

The task of determining a speaker's native language based only on his speeches in a second language is known as Native Language Identification or NLI. Due to its increasing applications in various domains of speech signal processing, this…

计算与语言 · 计算机科学 2018-11-15 Ahmed Nazim Uddin , Md Ashequr Rahman , Md. Rafidul Islam , Mohammad Ariful Haque

Hate detection has long been a challenging task for the NLP community. The task becomes complex in a code-mixed environment because the models must understand the context and the hate expressed through language alteration. Compared to the…

计算与语言 · 计算机科学 2024-10-22 Debajyoti Mazumder , Aakash Kumar , Jasabanta Patro

Despite an ever growing number of word representation models introduced for a large number of languages, there is a lack of a standardized technique to provide insights into what is captured by these models. Such insights would help the…

计算与语言 · 计算机科学 2019-12-12 Gözde Gül Şahin , Clara Vania , Ilia Kuznetsov , Iryna Gurevych

With nearly 1.5 billion people and more than 120 major languages, India represents one of the most diverse regions in the world. As multilingual Vision-Language Models (VLMs) gain prominence, robust evaluation methodologies are essential to…

The paper presents the submission of the team indicnlp@kgp to the EACL 2021 shared task "Offensive Language Identification in Dravidian Languages." The task aimed to classify different offensive content types in 3 code-mixed Dravidian…

计算与语言 · 计算机科学 2021-02-16 Kushal Kedia , Abhilash Nandy

This paper addresses challenges of Natural Language Processing (NLP) on non-canonical multilingual data in which two or more languages are mixed. It refers to code-switching which has become more popular in our daily life and therefore…

计算与语言 · 计算机科学 2016-10-10 Özlem Çetinoğlu , Sarah Schulz , Ngoc Thang Vu

While Large Language Models (LLMs) have significantly advanced Text-to-SQL performance, existing benchmarks predominantly focus on Western contexts and simplified schemas, leaving a gap in real-world, non-Western applications. We present…

计算与语言 · 计算机科学 2026-04-16 Aviral Dawar , Roshan Karanth , Vikram Goyal , Dhruv Kumar

Document layout analysis is essential for downstream tasks such as information retrieval, extraction, OCR, and digitization. However, existing large-scale datasets like PubLayNet and DocBank lack fine-grained region labels and multilingual…

计算机视觉与模式识别 · 计算机科学 2025-12-24 Oikantik Nath , Sahithi Kukkala , Mitesh Khapra , Ravi Kiran Sarvadevabhatla

Legal systems worldwide are inundated with exponential growth in cases and documents. There is an imminent need to develop NLP and ML techniques for automatically processing and understanding legal documents to streamline the legal system.…

计算与语言 · 计算机科学 2024-11-27 Abhinav Joshi , Shounak Paul , Akshat Sharma , Pawan Goyal , Saptarshi Ghosh , Ashutosh Modi

Language identification (LID) is an essential step in building high-quality multilingual datasets from web data. Existing LID tools (such as OpenLID or GlotLID) often struggle to identify closely related languages and to distinguish valid…

计算与语言 · 计算机科学 2026-02-24 Mariia Fedorova , Nikolay Arefyev , Maja Buljan , Jindřich Helcl , Stephan Oepen , Egil Rønningstad , Yves Scherrer