中文
相关论文

相关论文: UIT-HWDB: Using Transferring Method to Construct A…

200 篇论文

Scene-text image captioning requires fusing three information streams -- visual features, OCR-detected text, and linguistic knowledge -- to generate descriptions that faithfully integrate text visible in images. Existing fusion approaches…

计算机视觉与模式识别 · 计算机科学 2026-05-01 Nhi Ngoc-Yen Nguyen , Anh-Duc Nguyen , Nghia Hieu Nguyen , Kiet Van Nguyen , Ngan Luu-Thuy Nguyen

Training machines to synthesize diverse handwritings is an intriguing task. Recently, RNN-based methods have been proposed to generate stylized online Chinese characters. However, these methods mainly focus on capturing a person's overall…

计算机视觉与模式识别 · 计算机科学 2023-04-03 Gang Dai , Yifan Zhang , Qingfeng Wang , Qing Du , Zhuliang Yu , Zhuoman Liu , Shuangping Huang

Visual Question Answering (VQA) is a fundamental multimodal task that requires models to jointly understand visual and textual information. Early VQA systems relied heavily on language biases, motivating subsequent work to emphasize visual…

计算机视觉与模式识别 · 计算机科学 2026-03-12 Nguyen Anh Tuong , Phan Ba Duc , Nguyen Trung Quoc , Tran Dac Thinh , Dang Duy Lan , Nguyen Quoc Thinh , Tung Le

In the text classification problem, the imbalance of labels in datasets affect the performance of the text-classification models. Practically, the data about user comments on social networking sites not altogether appeared - the…

计算与语言 · 计算机科学 2020-10-12 Son T. Luu , Kiet Van Nguyen , Ngan Luu-Thuy Nguyen

We present ViT5, a pretrained Transformer-based encoder-decoder model for the Vietnamese language. With T5-style self-supervised pretraining, ViT5 is trained on a large corpus of high-quality and diverse Vietnamese texts. We benchmark ViT5…

计算与语言 · 计算机科学 2022-05-27 Long Phan , Hieu Tran , Hieu Nguyen , Trieu H. Trinh

Transliteration converts words in a source language (e.g., English) into words in a target language (e.g., Vietnamese). This conversion considers the phonological structure of the target language, as the transliterated output needs to be…

计算与语言 · 计算机科学 2019-02-21 Gia H. Ngo , Minh Nguyen , Nancy F. Chen

There are many difficulties facing a handwritten Arabic recognition system such as unlimited variation in human handwriting, similarities of distinct character shapes, interconnections of neighbouring characters and their position in the…

计算机视觉与模式识别 · 计算机科学 2014-02-27 Ahmed Sahlol , Cheng Suen

Despite the advent of deep learning in computer vision, the general handwriting recognition problem is far from solved. Most existing approaches focus on handwriting datasets that have clearly written text and carefully segmented labels. In…

计算机视觉与模式识别 · 计算机科学 2020-08-20 Hai Pham , Amrith Setlur , Saket Dingliwal , Tzu-Hsiang Lin , Barnabas Poczos , Kang Huang , Zhuo Li , Jae Lim , Collin McCormack , Tam Vu

In this paper, we demonstrate how a generative model can be used to build a better recognizer through the control of content and style. We are building an online handwriting recognizer from a modest amount of training samples. By training…

计算机视觉与模式识别 · 计算机科学 2021-10-15 Jen-Hao Rick Chang , Martin Bresler , Youssouf Chherawala , Adrien Delaye , Thomas Deselaers , Ryan Dixon , Oncel Tuzel

Image captioning is a crucial task with applications in a wide range of domains, including healthcare and education. Despite extensive research on English image captioning datasets, the availability of such datasets for Vietnamese remains…

计算机视觉与模式识别 · 计算机科学 2024-01-17 Anh-Cuong Pham , Van-Quang Nguyen , Thi-Hong Vuong , Quang-Thuy Ha

Converting written texts into their spoken forms is an essential problem in any text-to-speech (TTS) systems. However, building an effective text normalization solution for a real-world TTS system face two main challenges: (1) the semantic…

计算与语言 · 计算机科学 2022-09-08 Huu-Tien Dang , Thi-Hai-Yen Vuong , Xuan-Hieu Phan

Handwritten Text Recognition remains challenging due to the limited data, high writing style variance, and scripts with complex diacritics. Existing approaches, though partially address these issues, often struggle to generalize without…

计算机视觉与模式识别 · 计算机科学 2025-12-05 Pham Thach Thanh Truc , Dang Hoai Nam , Huynh Tong Dang Khoa , Vo Nguyen Le Duy

The advent of recurrent neural networks for handwriting recognition marked an important milestone reaching impressive recognition accuracies despite the great variability that we observe across different writing styles. Sequential…

计算机视觉与模式识别 · 计算机科学 2020-05-28 Lei Kang , Pau Riba , Marçal Rusiñol , Alicia Fornés , Mauricio Villegas

In this paper, we introduce Handwritten augmentation, a new data augmentation for handwritten character images. This method focuses on augmenting handwritten image data by altering the shape of input characters in training. The proposed…

计算机视觉与模式识别 · 计算机科学 2023-08-29 Mahendran N

Automatic Speech Recognition (ASR) performance is heavily dependent on the availability of large-scale, high-quality datasets. For low-resource languages, existing open-source ASR datasets often suffer from insufficient quality and…

计算与语言 · 计算机科学 2026-03-17 Thi Vu , Linh The Nguyen , Dat Quoc Nguyen

Identification of minimum number of local regions of a handwritten character image, containing well-defined discriminating features which are sufficient for a minimal but complete description of the character is a challenging task. A new…

计算机视觉与模式识别 · 计算机科学 2016-05-03 Ritesh Sarkhel , Amit K Saha , Nibaran Das

Developing effective scene text detection and recognition models hinges on extensive training data, which can be both laborious and costly to obtain, especially for low-resourced languages. Conventional methods tailored for Latin characters…

计算机视觉与模式识别 · 计算机科学 2024-10-25 Vannkinh Nom , Souhail Bakkali , Muhammad Muzzamil Luqman , Mickaël Coustaty , Jean-Marc Ogier

The reliance of humans over machines has never been so high such that from object classification in photographs to adding sound to silent movies everything can be performed with the help of deep learning and machine learning algorithms.…

计算机视觉与模式识别 · 计算机科学 2021-06-25 Samay Pashine , Ritik Dixit , Rishika Kushwah

Understanding signboard text in natural scenes is essential for real-world applications of Visual Question Answering (VQA), yet remains underexplored, particularly in low-resource languages. We introduce ViSignVQA, the first large-scale…

计算机视觉与模式识别 · 计算机科学 2025-12-30 Hieu Minh Nguyen , Tam Le-Thanh Dang , Kiet Van Nguyen

Vietnamese, the 20th most spoken language with over 102 million native speakers, lacks robust resources for key natural language processing tasks such as text segmentation and machine reading comprehension (MRC). To address this gap, we…

计算与语言 · 计算机科学 2025-06-23 Toan Nguyen Hai , Ha Nguyen Viet , Truong Quan Xuan , Duc Do Minh