English
Related papers

Related papers: Restoring Hebrew Diacritics Without a Dictionary

200 papers

Recent work attributes progress in NLP to large language models (LMs) with increased model size and large quantities of pretraining data. Despite this, current state-of-the-art LMs for Hebrew are both under-parameterized and under-trained…

Computation and Language · Computer Science 2022-12-20 Matan Eyal , Hila Noga , Roee Aharoni , Idan Szpektor , Reut Tsarfaty

Visual reasoning is challenging, requiring both precise object grounding and understanding complex spatial relationships. Existing methods fall into two camps: language-only chain-of-thought approaches, which demand large-scale (image,…

Computer Vision and Pattern Recognition · Computer Science 2025-12-10 Damiano Marsili , Georgia Gkioxari

Arabic handwriting is a consonantal and cursive writing. The analysis of Arabic script is further complicated due to obligatory dots/strokes that are placed above or below most letters and usually written delayed in order. Due to…

Computer Vision and Pattern Recognition · Computer Science 2015-10-20 Ibrahim Abdelaziz , Sherif Abdou , Hassanin Al-Barhamtoshy

Many information retrieval tasks require large labeled datasets for fine-tuning. However, such datasets are often unavailable, and their utility for real-world applications can diminish quickly due to domain shifts. To address this…

This study investigates the potential of Large Language Models (LLMs), particularly GPT-4o, for Optical Character Recognition (OCR) in low-resource scripts such as Urdu, Albanian, and Tajik, with English serving as a benchmark. Using a…

Machine Learning · Computer Science 2024-12-23 Muhammad Abdullah Sohail , Salaar Masood , Hamza Iqbal

Translating machine code into human-readable high-level languages is an open research problem in reverse engineering. Despite recent advancements in LLM-based decompilation to C, modern languages like Dart and Swift are unexplored. In this…

Software Engineering · Computer Science 2026-04-03 Raafat Abualazm , Ayman Abo Elhassan

We demonstrate a program that learns to pronounce Chinese text in Mandarin, without a pronunciation dictionary. From non-parallel streams of Chinese characters and Chinese pinyin syllables, it establishes a many-to-many mapping between…

Computation and Language · Computer Science 2020-10-13 Christopher Chu , Scot Fang , Kevin Knight

Large multilingual pretrained language models (mPLMs) have become the de facto state of the art for cross-lingual transfer in NLP. However, their large-scale deployment to many languages, besides pretraining data scarcity, is also hindered…

Computation and Language · Computer Science 2023-04-19 Sukannya Purkayastha , Sebastian Ruder , Jonas Pfeiffer , Iryna Gurevych , Ivan Vulić

Neural Machine Translation (NMT) on logographic source languages struggles when translating `unseen' characters, which never appear in the training data. One possible approach to this problem uses sub-character decomposition for training…

Computation and Language · Computer Science 2020-11-13 Danielle Saunders , Weston Feely , Bill Byrne

Previous studies have typically assumed that large language models are unable to accurately perform arithmetic operations, particularly multiplication of >8 digits, and operations involving decimals and fractions, without the use of…

Machine Learning · Computer Science 2023-09-13 Zhen Yang , Ming Ding , Qingsong Lv , Zhihuan Jiang , Zehai He , Yuyi Guo , Jinfeng Bai , Jie Tang

Vowels in Arabic are optional orthographic symbols written as diacritics above or below letters. In Arabic texts, typically more than 97 percent of written words do not explicitly show any of the vowels they contain; that is to say,…

Computation and Language · Computer Science 2019-05-13 Alexis Amid Neme , Sébastien Paumier

This paper presents a printed Bengali and English text OCR system developed by us using a single hidden BLSTM-CTC architecture having 128 units. Here, we did not use any peephole connection and dropout in the BLSTM, which helped us in…

Computer Vision and Pattern Recognition · Computer Science 2019-08-26 Debabrata Paul , Bidyut Baran Chaudhuri

An end-to-end, segmentation-free, deep learning model trained from scratch is proposed, leveraging DCNN for feature extraction, alongside Bidirectional Long-Short Term Memory (BLSTM) for sequence recognition and Connectionist Temporal…

Computer Vision and Pattern Recognition · Computer Science 2024-06-26 Sondos Aabed , Ahmad Khairaldin

We present a new pre-trained language model (PLM) for modern Hebrew, termed AlephBERTGimmel, which employs a much larger vocabulary (128K items) than standard Hebrew PLMs before. We perform a contrastive analysis of this model against all…

Computation and Language · Computer Science 2023-05-17 Eylon Gueta , Avi Shmidman , Shaltiel Shmidman , Cheyn Shmuel Shmidman , Joshua Guedalia , Moshe Koppel , Dan Bareket , Amit Seker , Reut Tsarfaty

Optical Character Recognition (OCR) for low-resource languages remains a significant challenge due to the scarcity of large-scale annotated training datasets. Languages such as Kashmiri, with approximately 7 million speakers and a complex…

Computation and Language · Computer Science 2026-01-23 Haq Nawaz Malik , Kh Mohmad Shafi , Tanveer Ahmad Reshi

Midrash collections are complex rabbinic works that consist of text in multiple languages, which evolved through long processes of unstable oral and written transmission. Determining the origin of a given passage in such a compilation is…

Computation and Language · Computer Science 2024-02-14 Shlomo Tannor , Nachum Dershowitz , Moshe Lavee

Code quality evaluation involves scoring generated code quality based on a reference code for a specific problem statement. Currently, there are two main forms of evaluating code quality: match-based evaluation and execution-based…

Software Engineering · Computer Science 2024-12-03 Fangzhou Xu , Sai Zhang , Zhenchang Xing , Xiaowang Zhang , Yahong Han , Zhiyong Feng

Recent work has shown that it is possible to train an $\textit{unsupervised}$ automatic speech recognition (ASR) system using only unpaired audio and text. Existing unsupervised ASR methods assume that no labeled data can be used for…

Audio and Speech Processing · Electrical Eng. & Systems 2024-02-19 Tatiana Likhomanenko , Loren Lugosch , Ronan Collobert

Epigraphy increasingly turns to modern artificial intelligence (AI) technologies such as machine learning (ML) for extracting insights from ancient inscriptions. However, scarce labeled data for training ML algorithms severely limits…

Computer Vision and Pattern Recognition · Computer Science 2023-10-12 Andrei C. Aioanei , Regine Hunziker-Rodewald , Konstantin Klein , Dominik L. Michels

Matching texts in highly inflected languages such as Arabic by simple stemming strategy is unlikely to perform well. In this paper, we present a strategy for automatic text matching technique for for inflectional languages, using Arabic as…

Computation and Language · Computer Science 2014-03-25 Tarek El-Shishtawy , Fatma El-Ghannam
‹ Prev 1 3 4 5 6 7 10 Next ›