English
Related papers

Related papers: Restoring Hebrew Diacritics Without a Dictionary

200 papers

Contrary to popular belief, Optical Character Recognition (OCR) remains a challenging problem when text occurs in unconstrained environments, like natural scenes, due to geometrical distortions, complex backgrounds, and diverse fonts. In…

Computer Vision and Pattern Recognition · Computer Science 2019-06-06 Marcin Namysl , Iuliu Konya

To remove redundant components of large language models (LLMs) without incurring significant computational costs, this work focuses on single-shot pruning without a retraining phase. We simplify the pruning process for Transformer-based…

Artificial Intelligence · Computer Science 2024-07-30 Jianwei Li , Yijun Dong , Qi Lei

Knowledge distillation is a widely adopted technique for transferring capabilities from LLMs to smaller, more efficient student models. However, unauthorized use of knowledge distillation takes unfair advantage of the considerable effort…

Artificial Intelligence · Computer Science 2026-04-20 Xinhang Ma , William Yeoh , Ning Zhang , Yevgeniy Vorobeychik

This paper introduces the L-ReLF (Low-Resource Lexical Framework), a novel, reproducible methodology for creating high-quality, structured lexical datasets for underserved languages. The lack of standardized terminology, exemplified by…

Computation and Language · Computer Science 2026-04-01 Anass Sedrati , Mounir Afifi , Reda Benkhadra

We propose Easymark, a family of embarrassingly simple yet effective watermarks. Text watermarking is becoming increasingly important with the advent of Large Language Models (LLM). LLMs can generate texts that cannot be distinguished from…

Machine Learning · Computer Science 2023-10-16 Ryoma Sato , Yuki Takezawa , Han Bao , Kenta Niwa , Makoto Yamada

Towards developing high-performing ASR for low-resource languages, approaches to address the lack of resources are to make use of data from multiple languages, and to augment the training data by creating acoustic variations. In this work…

Audio and Speech Processing · Electrical Eng. & Systems 2020-04-10 Chunxi Liu , Qiaochu Zhang , Xiaohui Zhang , Kritika Singh , Yatharth Saraf , Geoffrey Zweig

Despite their exceptional error-correcting properties, Reed-Solomon codes have been overlooked in distributed storage applications due to the common belief that they have poor repair bandwidth: A naive repair approach would require the…

Information Theory · Computer Science 2020-05-05 Hoang Dau , Iwan Duursma , Han Mao Kiah , Olgica Milenkovic

Recent embedding-based methods in unsupervised bilingual lexicon induction have shown good results, but generally have not leveraged orthographic (spelling) information, which can be helpful for pairs of related languages. This work…

Computation and Language · Computer Science 2020-02-04 Parker Riley , Daniel Gildea

We introduce the first unsupervised speech synthesis system based on a simple, yet effective recipe. The framework leverages recent work in unsupervised speech recognition as well as existing neural-based speech synthesis. Using only…

Sound · Computer Science 2022-04-21 Alexander H. Liu , Cheng-I Jeff Lai , Wei-Ning Hsu , Michael Auli , Alexei Baevski , James Glass

We present a novel inference scheme, self-speculative decoding, for accelerating Large Language Models (LLMs) without the need for an auxiliary model. This approach is characterized by a two-stage process: drafting and verification. The…

Computation and Language · Computer Science 2025-02-11 Jun Zhang , Jue Wang , Huan Li , Lidan Shou , Ke Chen , Gang Chen , Sharad Mehrotra

Traditional speech enhancement methods often oversimplify the task of restoration by focusing on a single type of distortion. Generative models that handle multiple distortions frequently struggle with phone reconstruction and…

Sound · Computer Science 2025-02-11 Tushar Dhyani , Florian Lux , Michele Mancusi , Giorgio Fabbro , Fritz Hohl , Ngoc Thang Vu

In this work, we present the development of a reverse transliteration model to convert romanized Malayalam to native script using an encoder-decoder framework built with attention-based bidirectional Long Short Term Memory (Bi-LSTM)…

Computation and Language · Computer Science 2024-12-16 Bajiyo Baiju , Kavya Manohar , Leena G Pillai , Elizabeth Sherly

Recent reference-based face restoration methods have received considerable attention due to their great capability in recovering high-frequency details on real low-quality images. However, most of these methods require a high-quality…

Computer Vision and Pattern Recognition · Computer Science 2020-08-04 Xiaoming Li , Chaofeng Chen , Shangchen Zhou , Xianhui Lin , Wangmeng Zuo , Lei Zhang

Proper nouns in Arabic Wikipedia are frequently undiacritized, creating ambiguity in pronunciation and interpretation, especially for transliterated named entities of foreign origin. While transliteration and diacritization have been…

Computation and Language · Computer Science 2025-06-24 Rawan Bondok , Mayar Nassar , Salam Khalifa , Kurt Micallef , Nizar Habash

In this paper, we address the problems of Arabic Text Classification and stemming using Transducers and Rational Kernels. We introduce a new stemming technique based on the use of Arabic patterns (Pattern Based Stemmer). Patterns are…

Computation and Language · Computer Science 2015-02-27 Attia Nehar , Djelloul Ziadi , Hadda Cherroun

The strength of obfuscated software has increased over the recent years. Compiler based obfuscation has become the de facto standard in the industry and recent papers also show that injection of obfuscation techniques is done at the…

Cryptography and Security · Computer Science 2019-09-17 Peter Garba , Matteo Favaro

Despite the existence of numerous Optical Character Recognition (OCR) tools, the lack of comprehensive open-source systems hampers the progress of document digitization in various low-resource languages, including Bengali. Low-resource…

We present crest, a tool for automatically proving (non-)confluence and termination of logically constrained rewrite systems. We compare crest to other tools for logically constrained rewriting. Extensive experiments demonstrate the promise…

Logic in Computer Science · Computer Science 2025-05-08 Jonas Schöpf , Aart Middeldorp

This study demonstrates that Large Language Models (LLMs) can transcribe historical handwritten documents with significantly higher accuracy than specialized Handwritten Text Recognition (HTR) software, while being faster and more…

Computer Vision and Pattern Recognition · Computer Science 2024-11-07 Mark Humphries , Lianne C. Leddy , Quinn Downton , Meredith Legace , John McConnell , Isabella Murray , Elizabeth Spence

In a global setting, texts contain transliterated names from many cultural origins. Correct transliteration depends not only on target and source languages but also, on the source language of the name. We introduce a novel methodology for…

Computation and Language · Computer Science 2019-11-28 Raphael Cohen , Michael Elhadad