中文
相关论文

相关论文: Script Normalization for Unconventional Writing of…

200 篇论文

Intelligent systems that aim at mastering language as humans do must deal with its semantic underspecification, namely, the possibility for a linguistic signal to convey only part of the information needed for communication to succeed.…

计算与语言 · 计算机科学 2023-06-09 Sandro Pezzelle

The performance of Language Models (LMs) on low-resource, morphologically rich languages like Sinhala remains largely unexplored, particularly regarding script variation in digital communication. Sinhala exhibits script duality, with…

计算与语言 · 计算机科学 2026-05-11 Minuri Rajapakse , Ruvan Weerasinghe

Recent years have seen a rise in interest for cross-lingual transfer between languages with similar typology, and between languages of various scripts. However, the interplay between language similarity and difference in script on…

计算与语言 · 计算机科学 2021-06-01 Samia Touileb , Jeremy Barnes

The presence of specific linguistic signals particular to a certain sub-group can become highly salient to language models during training. In automated decision-making settings, this may lead to biased outcomes when models rely on cues…

计算与语言 · 计算机科学 2025-09-05 Charmaine Barker , Dimitar Kazakov

Text normalization is a crucial technology for low-resource languages which lack rigid spelling conventions or that have undergone multiple spelling reforms. Low-resource text normalization has so far relied upon hand-crafted rules, which…

计算与语言 · 计算机科学 2023-12-25 Stefano Lusito , Edoardo Ferrante , Jean Maillard

In this paper, we define the task of gender rewriting in contexts involving two users (I and/or You) - first and second grammatical persons with independent grammatical gender preferences. We focus on Arabic, a gender-marking…

计算与语言 · 计算机科学 2022-05-05 Bashar Alhafni , Nizar Habash , Houda Bouamor

A large fraction of textual data available today contains various types of 'noise', such as OCR noise in digitized documents, noise due to informal writing style of users on microblogging sites, and so on. To enable tasks such as…

信息检索 · 计算机科学 2021-01-12 Anurag Roy , Shalmoli Ghosh , Kripabandhu Ghosh , Saptarshi Ghosh

The paper overviews the shared task on Real-Time Reverse Transliteration for Romanized Indo-Aryan languages. It focuses on the reverse transliteration of low-resourced languages in the Indo-Aryan family to their native scripts. Typing…

计算与语言 · 计算机科学 2025-02-25 Deshan Sumanathilaka , Isuri Anuradha , Ruvan Weerasinghe , Nicholas Micallef , Julian Hough

We describe an architecture for implementing spoken natural language dialogue interfaces to semi-autonomous systems, in which the central idea is to transform the input speech signal through successive levels of representation corresponding…

计算与语言 · 计算机科学 2007-05-23 Manny Rayner , Beth Ann Hockey , Frankie James

In recent years, automatic text summarization has witnessed significant advancement, particularly with the development of transformer-based models. However, the challenge of controlling the readability level of generated summaries remains…

计算与语言 · 计算机科学 2025-03-17 Mehmet Samet Duran , Tevfik Aytekin

With the rapid increase of transnational communication and cooperation, people frequently encounter multilingual scenarios in various situations. In this paper, we are concerned with a relatively new problem: script identification at word…

计算机视觉与模式识别 · 计算机科学 2015-05-13 Baoguang Shi , Cong Yao , Chengquan Zhang , Xiaowei Guo , Feiyue Huang , Xiang Bai

Handwriting recognition refers to the identification of written characters. Handwriting recognition has become an acute research area in recent years for the ease of access of computer science. In this paper primarily discussed On-line and…

计算机视觉与模式识别 · 计算机科学 2013-03-21 Dr. Firoj Parwej

Code-switching, the alternation of languages within a conversation or utterance, is a common communicative phenomenon that occurs in multilingual communities across the world. This survey reviews computational approaches for code-switched…

计算与语言 · 计算机科学 2020-07-24 Sunayana Sitaram , Khyathi Raghavi Chandu , Sai Krishna Rallabandi , Alan W Black

In this paper we address the scarcity of annotated data for NArabizi, a Romanized form of North African Arabic used mostly on social media, which poses challenges for Natural Language Processing (NLP). We introduce an enriched version of…

计算与语言 · 计算机科学 2024-12-06 Arij Riabi , Menel Mahamdi , Djamé Seddah

Signed languages are the primary means of communication for many deaf and hard of hearing individuals. Since signed languages exhibit all the fundamental linguistic properties of natural language, we believe that tools and theories of…

计算与语言 · 计算机科学 2021-07-26 Kayo Yin , Amit Moryossef , Julie Hochgesang , Yoav Goldberg , Malihe Alikhani

Due to reasons of convenience and lack of tech literacy, transliteration (i.e., Romanizing native scripts instead of using localization tools) is eminently prevalent in the context of low-resource languages such as Sinhala, which have their…

计算与语言 · 计算机科学 2025-03-05 Yomal De Mel , Kasun Wickramasinghe , Nisansa de Silva , Surangika Ranathunga

Spelling normalization for low resource languages is a challenging task because the patterns are hard to predict and large corpora are usually required to collect enough examples. This work shows a comparison of a neural model and character…

计算与语言 · 计算机科学 2020-10-21 Yiyuan Li , Antonios Anastasopoulos , Alan W Black

In general, speech processing models consist of a language model along with an acoustic model. Regardless of the language model's complexity and variants, three critical pre-processing steps are needed in language models: cleaning,…

音频与语音处理 · 电气工程与系统科学 2021-12-16 Romina Oji , Seyedeh Fatemeh Razavi , Sajjad Abdi Dehsorkh , Alireza Hariri , Hadi Asheri , Reshad Hosseini

A prototype system for the transliteration of diacritics-less Arabic manuscripts at the sub-word or part of Arabic word (PAW) level is developed. The system is able to read sub-words of the input manuscript using a set of skeleton-based…

计算机视觉与模式识别 · 计算机科学 2013-06-27 Reza Farrahi Moghaddam , Mohamed Cheriet , Thomas Milo , Robert Wisnovsky

Handwritten Text Recognition (HTR) under limited labeled data remains a challenging problem, particularly for Arabic-script languages. Although modern sequence-based recognizers perform well in high-resource settings, their accuracy…

计算机视觉与模式识别 · 计算机科学 2026-05-05 Sana Al-azzawi , Elisa Barney , Marcus Liwicki