中文
相关论文

相关论文: Punctuation restoration Model and Spacing Model fo…

200 篇论文

The performance of automatic speech recognition (ASR) models can be greatly improved by proper beam-search decoding with external language model (LM). There has been an increasing interest in Korean speech recognition, but not many studies…

计算与语言 · 计算机科学 2022-03-29 Kyuhong Shim , Hyewon Bae , Wonyong Sung

Comprehension of ancient texts plays an important role in archaeology and understanding of Chinese history and civilization. The rapid development of large language models needs benchmarks that can evaluate their comprehension of ancient…

计算与语言 · 计算机科学 2025-12-22 Zhihan Zhou , Daqian Shi , Rui Song , Lida Shi , Xiaolei Diao , Hao Xu

Good OCR results for historical printings rely on the availability of recognition models trained on diplomatic transcriptions as ground truth, which is both a scarce resource and time-consuming to generate. Instead of having to train a…

数字图书馆 · 计算机科学 2016-10-21 U. Springmann , F. Fink , K. U. Schulz

In digital images, the performance of optical aberration is a multivariate degradation, where the spectral of the scene, the lens imperfections, and the field of view together contribute to the results. Besides eliminating it at the…

计算机视觉与模式识别 · 计算机科学 2023-05-11 Shiqi Chen , Jinwen Zhou , Menghao Li , Yueting Chen , Tingting Jiang

Recently, much Chinese text error correction work has focused on Chinese Spelling Check (CSC) and Chinese Grammatical Error Diagnosis (CGED). In contrast, little attention has been paid to the complicated problem of Chinese Semantic Error…

计算与语言 · 计算机科学 2023-05-10 Bo Sun , Baoxin Wang , Yixuan Wang , Wanxiang Che , Dayong Wu , Shijin Wang , Ting Liu

Extracting fine-grained OCR text from aged documents in diacritic languages remains challenging due to unexpected artifacts, time-induced degradation, and lack of datasets. While standalone spell correction approaches have been proposed,…

计算与语言 · 计算机科学 2025-02-28 Thao Do , Dinh Phu Tran , An Vo , Daeyoung Kim

We describe a resource-based method of morphological annotation of written Korean text. Korean is an agglutinative language. The output of our system is a graph of morphemes annotated with accurate linguistic information. The language…

计算与语言 · 计算机科学 2007-11-22 Hyun-Gue Huh , Eric Laporte

Punctuation restoration is a crucial step after Automatic Speech Recognition (ASR) systems to enhance transcript readability and facilitate subsequent NLP tasks. Nevertheless, conventional lexical-based approaches are inadequate for solving…

计算与语言 · 计算机科学 2024-02-07 Xiliang Zhu , Chia-Tien Chang , Shayna Gardiner , David Rossouw , Jonas Robertson

Chinese Spelling Correction (CSC) aims to detect and correct spelling errors in Chinese sentences caused by phonetic or visual similarities. While current CSC models integrate pinyin or glyph features and have shown significant…

计算与语言 · 计算机科学 2024-09-10 Lei Sheng , Shuai-Shuai Xu

Automatic speech recognition (ASR) systems in the medical domain that focus on transcribing clinical dictations and doctor-patient conversations often pose many challenges due to the complexity of the domain. ASR output typically undergoes…

计算与语言 · 计算机科学 2020-07-14 Monica Sunkara , Srikanth Ronanki , Kalpit Dixit , Sravan Bodapati , Katrin Kirchhoff

With the recent advance in neural machine translation demonstrating its importance, research on quality estimation (QE) has been steadily progressing. QE aims to automatically predict the quality of machine translation (MT) output without…

计算与语言 · 计算机科学 2022-11-30 Sugyeong Eo , Chanjun Park , Hyeonseok Moon , Jaehyung Seo , Gyeongmin Kim , Jungseob Lee , Heuiseok Lim

Ancient people translated classical Chinese into Japanese using a system of annotations placed around characters. We abstract this process as sequence tagging tasks and fit them into modern language technologies. The research on this…

计算与语言 · 计算机科学 2026-01-22 Zilong Li , Jie Cao

For readability and disambiguation of the written text, appropriate word segmentation is recommended for documentation, and it also holds for the digitized texts. If the language is agglutinative while far from scriptio continua, for…

计算与语言 · 计算机科学 2021-05-05 Won Ik Cho , Sung Jun Cheon , Woo Hyun Kang , Ji Won Kim , Nam Soo Kim

Generative models, widely utilized in various applications, can often struggle with prompts corresponding to partial tokens. This struggle stems from tokenization, where partial tokens fall out of distribution during inference, leading to…

While Korean historical documents are invaluable cultural heritage, understanding those documents requires in-depth Hanja expertise. Hanja is an ancient language used in Korea before the 20th century, whose characters were borrowed from old…

计算与语言 · 计算机科学 2025-01-22 Seyoung Song , Haneul Yoo , Jiho Jin , Kyunghyun Cho , Alice Oh

In recent years, after the neural-network-based method was proposed, the accuracy of the Chinese word segmentation task has made great progress. However, when dealing with out-of-vocabulary words, there is still a large error rate. We used…

计算与语言 · 计算机科学 2019-01-18 Yung-Sung Chuang

Large language models have made significant advancements in various natural language processing tasks, including coreference resolution. However, traditional methods often fall short in effectively distinguishing referential relationships…

计算与语言 · 计算机科学 2025-04-09 Xingzu Liu , Songhang deng , Mingbang Wang , Zhang Dong , Le Dai , Jiyuan Li , Ruilin Nong

This paper describes the algorithm for translating English negative sentences into Korean in English-Korean Machine Translation (EKMT). The proposed algorithm is based on the comparative study of English and Korean negative sentences. The…

计算与语言 · 计算机科学 2015-12-29 Chung-Hyok Jang , Kwang-Hyok Kim

Sentence Simplification is a valuable technique that can benefit language learners and children a lot. However, current research focuses more on English sentence simplification. The development of Chinese sentence simplification is…

计算与语言 · 计算机科学 2023-06-08 Shiping Yang , Renliang Sun , Xiaojun Wan

Kurdish libraries have many historical publications that were printed back in the early days when printing devices were brought to Kurdistan. Having a good Optical Character Recognition (OCR) to help process these publications and…

计算与语言 · 计算机科学 2024-04-10 Blnd Yaseen , Hossein Hassani