中文
相关论文

相关论文: Restoring Hebrew Diacritics Without a Dictionary

200 篇论文

Contrary to popular belief, Optical Character Recognition (OCR) remains a challenging problem when text occurs in unconstrained environments, like natural scenes, due to geometrical distortions, complex backgrounds, and diverse fonts. In…

计算机视觉与模式识别 · 计算机科学 2019-06-06 Marcin Namysl , Iuliu Konya

To remove redundant components of large language models (LLMs) without incurring significant computational costs, this work focuses on single-shot pruning without a retraining phase. We simplify the pruning process for Transformer-based…

人工智能 · 计算机科学 2024-07-30 Jianwei Li , Yijun Dong , Qi Lei

Knowledge distillation is a widely adopted technique for transferring capabilities from LLMs to smaller, more efficient student models. However, unauthorized use of knowledge distillation takes unfair advantage of the considerable effort…

人工智能 · 计算机科学 2026-04-20 Xinhang Ma , William Yeoh , Ning Zhang , Yevgeniy Vorobeychik

This paper introduces the L-ReLF (Low-Resource Lexical Framework), a novel, reproducible methodology for creating high-quality, structured lexical datasets for underserved languages. The lack of standardized terminology, exemplified by…

计算与语言 · 计算机科学 2026-04-01 Anass Sedrati , Mounir Afifi , Reda Benkhadra

We propose Easymark, a family of embarrassingly simple yet effective watermarks. Text watermarking is becoming increasingly important with the advent of Large Language Models (LLM). LLMs can generate texts that cannot be distinguished from…

机器学习 · 计算机科学 2023-10-16 Ryoma Sato , Yuki Takezawa , Han Bao , Kenta Niwa , Makoto Yamada

Towards developing high-performing ASR for low-resource languages, approaches to address the lack of resources are to make use of data from multiple languages, and to augment the training data by creating acoustic variations. In this work…

音频与语音处理 · 电气工程与系统科学 2020-04-10 Chunxi Liu , Qiaochu Zhang , Xiaohui Zhang , Kritika Singh , Yatharth Saraf , Geoffrey Zweig

Despite their exceptional error-correcting properties, Reed-Solomon codes have been overlooked in distributed storage applications due to the common belief that they have poor repair bandwidth: A naive repair approach would require the…

信息论 · 计算机科学 2020-05-05 Hoang Dau , Iwan Duursma , Han Mao Kiah , Olgica Milenkovic

Recent embedding-based methods in unsupervised bilingual lexicon induction have shown good results, but generally have not leveraged orthographic (spelling) information, which can be helpful for pairs of related languages. This work…

计算与语言 · 计算机科学 2020-02-04 Parker Riley , Daniel Gildea

We introduce the first unsupervised speech synthesis system based on a simple, yet effective recipe. The framework leverages recent work in unsupervised speech recognition as well as existing neural-based speech synthesis. Using only…

声音 · 计算机科学 2022-04-21 Alexander H. Liu , Cheng-I Jeff Lai , Wei-Ning Hsu , Michael Auli , Alexei Baevski , James Glass

We present a novel inference scheme, self-speculative decoding, for accelerating Large Language Models (LLMs) without the need for an auxiliary model. This approach is characterized by a two-stage process: drafting and verification. The…

计算与语言 · 计算机科学 2025-02-11 Jun Zhang , Jue Wang , Huan Li , Lidan Shou , Ke Chen , Gang Chen , Sharad Mehrotra

Traditional speech enhancement methods often oversimplify the task of restoration by focusing on a single type of distortion. Generative models that handle multiple distortions frequently struggle with phone reconstruction and…

声音 · 计算机科学 2025-02-11 Tushar Dhyani , Florian Lux , Michele Mancusi , Giorgio Fabbro , Fritz Hohl , Ngoc Thang Vu

In this work, we present the development of a reverse transliteration model to convert romanized Malayalam to native script using an encoder-decoder framework built with attention-based bidirectional Long Short Term Memory (Bi-LSTM)…

计算与语言 · 计算机科学 2024-12-16 Bajiyo Baiju , Kavya Manohar , Leena G Pillai , Elizabeth Sherly

Recent reference-based face restoration methods have received considerable attention due to their great capability in recovering high-frequency details on real low-quality images. However, most of these methods require a high-quality…

计算机视觉与模式识别 · 计算机科学 2020-08-04 Xiaoming Li , Chaofeng Chen , Shangchen Zhou , Xianhui Lin , Wangmeng Zuo , Lei Zhang

Proper nouns in Arabic Wikipedia are frequently undiacritized, creating ambiguity in pronunciation and interpretation, especially for transliterated named entities of foreign origin. While transliteration and diacritization have been…

计算与语言 · 计算机科学 2025-06-24 Rawan Bondok , Mayar Nassar , Salam Khalifa , Kurt Micallef , Nizar Habash

In this paper, we address the problems of Arabic Text Classification and stemming using Transducers and Rational Kernels. We introduce a new stemming technique based on the use of Arabic patterns (Pattern Based Stemmer). Patterns are…

计算与语言 · 计算机科学 2015-02-27 Attia Nehar , Djelloul Ziadi , Hadda Cherroun

The strength of obfuscated software has increased over the recent years. Compiler based obfuscation has become the de facto standard in the industry and recent papers also show that injection of obfuscation techniques is done at the…

密码学与安全 · 计算机科学 2019-09-17 Peter Garba , Matteo Favaro

Despite the existence of numerous Optical Character Recognition (OCR) tools, the lack of comprehensive open-source systems hampers the progress of document digitization in various low-resource languages, including Bengali. Low-resource…

We present crest, a tool for automatically proving (non-)confluence and termination of logically constrained rewrite systems. We compare crest to other tools for logically constrained rewriting. Extensive experiments demonstrate the promise…

计算机科学中的逻辑 · 计算机科学 2025-05-08 Jonas Schöpf , Aart Middeldorp

This study demonstrates that Large Language Models (LLMs) can transcribe historical handwritten documents with significantly higher accuracy than specialized Handwritten Text Recognition (HTR) software, while being faster and more…

计算机视觉与模式识别 · 计算机科学 2024-11-07 Mark Humphries , Lianne C. Leddy , Quinn Downton , Meredith Legace , John McConnell , Isabella Murray , Elizabeth Spence

In a global setting, texts contain transliterated names from many cultural origins. Correct transliteration depends not only on target and source languages but also, on the source language of the name. We introduce a novel methodology for…

计算与语言 · 计算机科学 2019-11-28 Raphael Cohen , Michael Elhadad