中文
相关论文

相关论文: Processing South Asian Languages Written in the La…

200 篇论文

Text simplification is a valuable technique. However, current research is limited to sentence simplification. In this paper, we define and investigate a new task of document-level text simplification, which aims to simplify a document…

计算与语言 · 计算机科学 2021-10-12 Renliang Sun , Hanqi Jin , Xiaojun Wan

The wide accessibility of social media has provided linguistically under-represented communities with an extraordinary opportunity to create content in their native languages. This, however, comes with certain challenges in script…

计算与语言 · 计算机科学 2023-05-29 Sina Ahmadi , Antonios Anastasopoulos

Sentiment Analysis (SA) is an action research area in the digital age. With rapid and constant growth of online social media sites and services, and the increasing amount of textual data such as - statuses, comments, reviews etc. available…

计算与语言 · 计算机科学 2016-11-28 A. Hassan , M. R. Amin , N. Mohammed , A. K. A. Azad

Automatic language identification is a natural language processing problem that tries to determine the natural language of a given content. In this paper we present a statistical method for automatic language identification of written text…

计算与语言 · 计算机科学 2018-06-15 Ciprian-Octavian Truică , Julien Velcin , Alexandru Boicea

In this paper, we introduce a data-driven approach to transliterating Uzbek dictionary words from the Cyrillic script into the Latin script, and vice versa. We heuristically align characters of words in the source script with sub-strings of…

计算与语言 · 计算机科学 2021-01-14 B. Mansurov , A. Mansurov

We release S\={a}mayik, a dataset of around 53,000 parallel English-Sanskrit sentences, written in contemporary prose. Sanskrit is a classical language still in sustenance and has a rich documented heritage. However, due to the limited…

This paper presents a novel task of extracting low-resourced and noisy Latin fragments from mixed-language historical documents with varied layouts. We benchmark and evaluate the performance of large foundation models against a multimodal…

计算与语言 · 计算机科学 2026-02-09 Yu Wu , Ke Shu , Jonas Fischer , Lidia Pivovarova , David Rosson , Eetu Mäkelä , Mikko Tolonen

This paper focuses on developing translation models and related applications for 36 Indian languages, including Assamese, Awadhi, Bengali, Bhojpuri, Braj, Bodo, Dogri, English, Konkani, Gondi, Gujarati, Hindi, Hinglish, Ho, Kannada, Kangri,…

计算与语言 · 计算机科学 2025-01-03 Vandan Mujadia , Dipti Misra Sharma

We introduce Stanza, an open-source Python natural language processing toolkit supporting 66 human languages. Compared to existing widely used toolkits, Stanza features a language-agnostic fully neural pipeline for text analysis, including…

计算与语言 · 计算机科学 2020-04-24 Peng Qi , Yuhao Zhang , Yuhui Zhang , Jason Bolton , Christopher D. Manning

The digitisation of classical Sanskrit literature is impeded by a scarcity of annotated resources, particularly for Named Entity Recognition. While recent methodologies utilise generic Large Language Models (LLMs) for data augmentation,…

计算与语言 · 计算机科学 2026-04-30 Akhil Rajeev P , Annarao Kulkarni

Scene-text recognition is remarkably better in Latin languages than the non-Latin languages due to several factors like multiple fonts, simplistic vocabulary statistics, updated data generation tools, and writing systems. This paper…

计算机视觉与模式识别 · 计算机科学 2022-01-11 Sanjana Gunna , Rohit Saluja , C. V. Jawahar

Reading scene text, that is, text appearing in images, has numerous application areas, including assistive technology, search, and e-commerce. Although scene text recognition in English has advanced significantly and is often considered…

计算机视觉与模式识别 · 计算机科学 2026-04-13 Anik De , Abhirama Subramanyam Penamakuri , Rajeev Yadav , Aditya Rathore , Harshiv Shah , Devesh Sharma , Sagar Agarwal , Pravin Kumar , Anand Mishra

This article presents a Bangla handwriting dataset named BanglaWriting that contains single-page handwritings of 260 individuals of different personalities and ages. Each page includes bounding-boxes that bounds each word, along with the…

计算机视觉与模式识别 · 计算机科学 2022-08-22 M. F. Mridha , Abu Quwsar Ohi , M. Ameer Ali , Mazedul Islam Emon , Muhammad Mohsin Kabir

This paper describes the submissions by team HWR to the Dravidian Language Identification (DLI) shared task organized at VarDial 2021 workshop. The DLI training set includes 16,674 YouTube comments written in Roman script containing…

计算与语言 · 计算机科学 2021-03-10 Tommi Jauhiainen , Tharindu Ranasinghe , Marcos Zampieri

This paper describes the development of a multilingual, manually annotated dataset for three under-resourced Dravidian languages generated from social media comments. The dataset was annotated for sentiment analysis and offensive language…

With the rapid increase of transnational communication and cooperation, people frequently encounter multilingual scenarios in various situations. In this paper, we are concerned with a relatively new problem: script identification at word…

计算机视觉与模式识别 · 计算机科学 2015-05-13 Baoguang Shi , Cong Yao , Chengquan Zhang , Xiaowei Guo , Feiyue Huang , Xiang Bai

Machine Transliteration provides the ability to transliterate a basic language into different languages in a computational way. Transliteration is an important technical process that has caught the attention most recently. The Sinhala…

计算与语言 · 计算机科学 2024-04-23 Maneesha U. Athukorala , Deshan K. Sumanathilaka

We present Latin BERT, a contextual language model for the Latin language, trained on 642.7 million words from a variety of sources spanning the Classical era to the 21st century. In a series of case studies, we illustrate the affordances…

计算与语言 · 计算机科学 2020-09-22 David Bamman , Patrick J. Burns

Traditional approaches to semantic parsing (SP) work by training individual models for each available parallel dataset of text-meaning pairs. In this paper, we explore the idea of polyglot semantic translation, or learning semantic parsing…

计算与语言 · 计算机科学 2018-07-12 Kyle Richardson , Jonathan Berant , Jonas Kuhn

Data-driven approaches for dependency parsing have been of great interest in Natural Language Processing for the past couple of decades. However, Sanskrit still lacks a robust purely data-driven dependency parser, probably with an exception…

计算与语言 · 计算机科学 2020-04-20 Amrith Krishna , Ashim Gupta , Deepak Garasangi , Jivnesh Sandhan , Pavankumar Satuluri , Pawan Goyal