中文
相关论文

相关论文: Restoring Hebrew Diacritics Without a Dictionary

200 篇论文

Recent work attributes progress in NLP to large language models (LMs) with increased model size and large quantities of pretraining data. Despite this, current state-of-the-art LMs for Hebrew are both under-parameterized and under-trained…

计算与语言 · 计算机科学 2022-12-20 Matan Eyal , Hila Noga , Roee Aharoni , Idan Szpektor , Reut Tsarfaty

Visual reasoning is challenging, requiring both precise object grounding and understanding complex spatial relationships. Existing methods fall into two camps: language-only chain-of-thought approaches, which demand large-scale (image,…

计算机视觉与模式识别 · 计算机科学 2025-12-10 Damiano Marsili , Georgia Gkioxari

Arabic handwriting is a consonantal and cursive writing. The analysis of Arabic script is further complicated due to obligatory dots/strokes that are placed above or below most letters and usually written delayed in order. Due to…

计算机视觉与模式识别 · 计算机科学 2015-10-20 Ibrahim Abdelaziz , Sherif Abdou , Hassanin Al-Barhamtoshy

Many information retrieval tasks require large labeled datasets for fine-tuning. However, such datasets are often unavailable, and their utility for real-world applications can diminish quickly due to domain shifts. To address this…

This study investigates the potential of Large Language Models (LLMs), particularly GPT-4o, for Optical Character Recognition (OCR) in low-resource scripts such as Urdu, Albanian, and Tajik, with English serving as a benchmark. Using a…

机器学习 · 计算机科学 2024-12-23 Muhammad Abdullah Sohail , Salaar Masood , Hamza Iqbal

Translating machine code into human-readable high-level languages is an open research problem in reverse engineering. Despite recent advancements in LLM-based decompilation to C, modern languages like Dart and Swift are unexplored. In this…

软件工程 · 计算机科学 2026-04-03 Raafat Abualazm , Ayman Abo Elhassan

We demonstrate a program that learns to pronounce Chinese text in Mandarin, without a pronunciation dictionary. From non-parallel streams of Chinese characters and Chinese pinyin syllables, it establishes a many-to-many mapping between…

计算与语言 · 计算机科学 2020-10-13 Christopher Chu , Scot Fang , Kevin Knight

Large multilingual pretrained language models (mPLMs) have become the de facto state of the art for cross-lingual transfer in NLP. However, their large-scale deployment to many languages, besides pretraining data scarcity, is also hindered…

计算与语言 · 计算机科学 2023-04-19 Sukannya Purkayastha , Sebastian Ruder , Jonas Pfeiffer , Iryna Gurevych , Ivan Vulić

Neural Machine Translation (NMT) on logographic source languages struggles when translating `unseen' characters, which never appear in the training data. One possible approach to this problem uses sub-character decomposition for training…

计算与语言 · 计算机科学 2020-11-13 Danielle Saunders , Weston Feely , Bill Byrne

Previous studies have typically assumed that large language models are unable to accurately perform arithmetic operations, particularly multiplication of >8 digits, and operations involving decimals and fractions, without the use of…

机器学习 · 计算机科学 2023-09-13 Zhen Yang , Ming Ding , Qingsong Lv , Zhihuan Jiang , Zehai He , Yuyi Guo , Jinfeng Bai , Jie Tang

Vowels in Arabic are optional orthographic symbols written as diacritics above or below letters. In Arabic texts, typically more than 97 percent of written words do not explicitly show any of the vowels they contain; that is to say,…

计算与语言 · 计算机科学 2019-05-13 Alexis Amid Neme , Sébastien Paumier

This paper presents a printed Bengali and English text OCR system developed by us using a single hidden BLSTM-CTC architecture having 128 units. Here, we did not use any peephole connection and dropout in the BLSTM, which helped us in…

计算机视觉与模式识别 · 计算机科学 2019-08-26 Debabrata Paul , Bidyut Baran Chaudhuri

An end-to-end, segmentation-free, deep learning model trained from scratch is proposed, leveraging DCNN for feature extraction, alongside Bidirectional Long-Short Term Memory (BLSTM) for sequence recognition and Connectionist Temporal…

计算机视觉与模式识别 · 计算机科学 2024-06-26 Sondos Aabed , Ahmad Khairaldin

We present a new pre-trained language model (PLM) for modern Hebrew, termed AlephBERTGimmel, which employs a much larger vocabulary (128K items) than standard Hebrew PLMs before. We perform a contrastive analysis of this model against all…

Optical Character Recognition (OCR) for low-resource languages remains a significant challenge due to the scarcity of large-scale annotated training datasets. Languages such as Kashmiri, with approximately 7 million speakers and a complex…

计算与语言 · 计算机科学 2026-01-23 Haq Nawaz Malik , Kh Mohmad Shafi , Tanveer Ahmad Reshi

Midrash collections are complex rabbinic works that consist of text in multiple languages, which evolved through long processes of unstable oral and written transmission. Determining the origin of a given passage in such a compilation is…

计算与语言 · 计算机科学 2024-02-14 Shlomo Tannor , Nachum Dershowitz , Moshe Lavee

Code quality evaluation involves scoring generated code quality based on a reference code for a specific problem statement. Currently, there are two main forms of evaluating code quality: match-based evaluation and execution-based…

软件工程 · 计算机科学 2024-12-03 Fangzhou Xu , Sai Zhang , Zhenchang Xing , Xiaowang Zhang , Yahong Han , Zhiyong Feng

Recent work has shown that it is possible to train an $\textit{unsupervised}$ automatic speech recognition (ASR) system using only unpaired audio and text. Existing unsupervised ASR methods assume that no labeled data can be used for…

音频与语音处理 · 电气工程与系统科学 2024-02-19 Tatiana Likhomanenko , Loren Lugosch , Ronan Collobert

Epigraphy increasingly turns to modern artificial intelligence (AI) technologies such as machine learning (ML) for extracting insights from ancient inscriptions. However, scarce labeled data for training ML algorithms severely limits…

计算机视觉与模式识别 · 计算机科学 2023-10-12 Andrei C. Aioanei , Regine Hunziker-Rodewald , Konstantin Klein , Dominik L. Michels

Matching texts in highly inflected languages such as Arabic by simple stemming strategy is unlikely to perform well. In this paper, we present a strategy for automatic text matching technique for for inflectional languages, using Arabic as…

计算与语言 · 计算机科学 2014-03-25 Tarek El-Shishtawy , Fatma El-Ghannam