中文
相关论文

相关论文: Script Normalization for Unconventional Writing of…

200 篇论文

As large language models (LLMs) are trained on increasingly diverse and extensive multilingual corpora, they demonstrate cross-lingual transfer capabilities. However, these capabilities often fail to effectively extend to low-resource…

计算与语言 · 计算机科学 2025-09-23 Wenhao Zhuang , Yuan Sun , Xiaobing Zhao

Whether it is in the form of transcribed conversations, blog posts, or tweets, qualitative data provides a reader with rich insight into both the overarching trends as well as the diversity of human ideas expressed through text. Handling…

人机交互 · 计算机科学 2022-09-27 Huyen N. Nguyen , Tommy Dang , Kathleen A. Bowe

Our goal is to create a convenient natural language interface for performing well-specified but complex actions such as analyzing data, manipulating text, and querying databases. However, existing natural language interfaces for such tasks…

计算与语言 · 计算机科学 2017-04-25 Sida I. Wang , Samuel Ginn , Percy Liang , Christoper D. Manning

Sign languages, of any geographical or accentual variation, understandably face continuous scrutiny under the ever present popularity of verbal dictation and audism. Through this, many potential problems arise with the current lack of…

人工智能 · 计算机科学 2026-05-01 Nina Seron-Abouelfadil , Poppy Fynes

The use of short text messages in social media and instant messaging has become a popular communication channel during the last years. This rising popularity has caused an increment in messaging threats such as spam, phishing or malware as…

计算与语言 · 计算机科学 2014-08-19 Alejandro Mosquera , Lamine Aouad , Slawomir Grzonkowski , Dylan Morss

Subwords are the most widely used output units in end-to-end speech recognition. They combine the best of two worlds by modeling the majority of frequent words directly and at the same time allow open vocabulary speech recognition by…

音频与语音处理 · 电气工程与系统科学 2020-08-11 Egor Lakomkin , Jahn Heymann , Ilya Sklyar , Simon Wiesler

We observe a recent behaviour on social media, in which users intentionally remove consonantal dots from Arabic letters, in order to bypass content-classification algorithms. Content classification is typically done by fine-tuning…

计算与语言 · 计算机科学 2021-11-19 Aviad Rom , Kfir Bar

Despite speaking mutually intelligible varieties of the same language, speakers of Tajik Persian, written in a modified Cyrillic alphabet, cannot read Iranian and Afghan texts written in the Perso-Arabic script. As the vast majority of…

计算与语言 · 计算机科学 2025-02-28 Rayyan Merchant , Akhilesh Kakolu Ramarao , Kevin Tang

User generated text on social media often suffers from a lot of undesired characteristics including hatespeech, abusive language, insults etc. that are targeted to attack or abuse a specific group of people. Often such text is written…

计算与语言 · 计算机科学 2019-10-03 Sravan Babu Bodapati , Spandana Gella , Kasturi Bhattacharjee , Yaser Al-Onaizan

Linguistic bias in online news and social media is widespread but difficult to measure. Yet, its identification and quantification remain difficult due to subjectivity, context dependence, and the scarcity of high-quality gold-label…

信息检索 · 计算机科学 2025-12-17 Fabian Haak , Philipp Schaer

In today's software world with its cornucopia of reusable software libraries, when a programmer is faced with a programming task that they suspect can be completed through the use of a library, they often look for code examples using a…

软件工程 · 计算机科学 2021-10-08 Geert Heyman , Rafael Huysegems , Pascal Justen , Tom Van Cutsem

A large number of annotated training images is crucial for training successful scene text recognition models. However, collecting sufficient datasets can be a labor-intensive and costly process, particularly for low-resource languages. To…

计算机视觉与模式识别 · 计算机科学 2023-06-28 Yangchen Xie , Xinyuan Chen , Hongjian Zhan , Palaiahankote Shivakum , Bing Yin , Cong Liu , Yue Lu

Recent embedding-based methods in unsupervised bilingual lexicon induction have shown good results, but generally have not leveraged orthographic (spelling) information, which can be helpful for pairs of related languages. This work…

计算与语言 · 计算机科学 2020-02-04 Parker Riley , Daniel Gildea

Large, pre-trained representation models trained using self-supervised learning have gained popularity in various fields of machine learning because they are able to extract high-quality salient features from input data. As such, they have…

音频与语音处理 · 电气工程与系统科学 2023-06-16 Hejung Yang , Hong-Goo Kang

Code summarization generates brief natural language descriptions of source code pieces, which can assist developers in understanding code and reduce documentation workload. Recent neural models on code summarization are trained and…

软件工程 · 计算机科学 2022-10-24 Rui Xie , Tianxiang Hu , Wei Ye , Shikun Zhang

Social Media platforms have been seeing adoption and growth in their usage over time. This growth has been further accelerated with the lockdown in the past year when people's interaction, conversation, and expression were limited…

计算与语言 · 计算机科学 2022-04-06 Ekagra Ranjan , Naman Poddar

Script knowledge plays a central role in text understanding and is relevant for a variety of downstream tasks. In this paper, we consider two recent datasets which provide a rich and general representation of script events in terms of…

计算与语言 · 计算机科学 2019-05-21 Simon Ostermann , Michael Roth , Stefan Thater , Manfred Pinkal

Historic variations of spelling poses a challenge for full-text search or natural language processing on historical digitized texts. To minimize the gap between the historic orthography and contemporary spelling, usually an automatic…

计算与语言 · 计算机科学 2025-02-26 Anton Ehrmanntraut

There are many difficulties facing a handwritten Arabic recognition system such as unlimited variation in human handwriting, similarities of distinct character shapes, interconnections of neighbouring characters and their position in the…

计算机视觉与模式识别 · 计算机科学 2014-02-27 Ahmed Sahlol , Cheng Suen

As a digraphic language, the Persian language utilizes two written standards: Perso-Arabic in Afghanistan and Iran, and Tajik-Cyrillic in Tajikistan. Despite the significant similarity between the dialects of each country, script…

计算与语言 · 计算机科学 2025-10-10 Rayyan Merchant , Kevin Tang
‹ 上一页 1 8 9 10 下一页 ›