中文
相关论文

相关论文: STELLAR: Scene Text Editor for Low-Resource Langua…

200 篇论文

Text line detection is crucial for any application associated with Automatic Text Recognition or Keyword Spotting. Modern algorithms perform good on well-established datasets since they either comprise clean data or simple/homogeneous page…

计算机视觉与模式识别 · 计算机科学 2017-12-12 Tobias Grüning , Roger Labahn , Markus Diem , Florian Kleber , Stefan Fiel

The paper proposes a new text recognition network for scene-text images. Many state-of-the-art methods employ the attention mechanism either in the text encoder or decoder for the text alignment. Although the encoder-based attention yields…

计算机视觉与模式识别 · 计算机科学 2021-04-27 Usman Sajid , Michael Chow , Jin Zhang , Taejoon Kim , Guanghui Wang

Text-to-image generative models have garnered immense attention for their ability to produce high-fidelity images from text prompts. Among these, Stable Diffusion distinguishes itself as a leading open-source model in this fast-growing…

计算机视觉与模式识别 · 计算机科学 2024-03-12 Shih-Ying Yeh , Yu-Guan Hsieh , Zhidong Gao , Bernard B W Yang , Giyeong Oh , Yanmin Gong

This paper aims to re-assess scene text recognition (STR) from a data-oriented perspective. We begin by revisiting the six commonly used benchmarks in STR and observe a trend of performance saturation, whereby only 2.91% of the benchmark…

计算机视觉与模式识别 · 计算机科学 2023-07-20 Qing Jiang , Jiapeng Wang , Dezhi Peng , Chongyu Liu , Lianwen Jin

We propose an unsupervised adaptation framework, Self-TAught Recognizer (STAR), which leverages unlabeled data to enhance the robustness of automatic speech recognition (ASR) systems in diverse target domains, such as noise and accents.…

计算与语言 · 计算机科学 2024-05-24 Yuchen Hu , Chen Chen , Chao-Han Huck Yang , Chengwei Qin , Pin-Yu Chen , Eng Siong Chng , Chao Zhang

Multimodal image-tabular learning is gaining attention, yet it faces challenges due to limited labeled data. While earlier work has applied self-supervised learning (SSL) to unlabeled data, its task-agnostic nature often results in learning…

计算机视觉与模式识别 · 计算机科学 2025-03-18 Siyi Du , Xinzhe Luo , Declan P. O'Regan , Chen Qin

The ability to fine-tune generative models for text-to-image generation tasks is crucial, particularly facing the complexity involved in accurately interpreting and visualizing textual inputs. While LoRA is efficient for language model…

计算机视觉与模式识别 · 计算机科学 2024-05-13 Mohan Zhou , Yalong Bai , Qing Yang , Tiejun Zhao

Text-to-LiDAR generation can customize 3D data with rich structures and diverse scenes for downstream tasks. However, the scarcity of Text-LiDAR pairs often causes insufficient training priors, generating overly smooth 3D scenes. Moreover,…

计算机视觉与模式识别 · 计算机科学 2025-12-16 Wentao Qu , Guofeng Mei , Yang Wu , Yongshun Gong , Xiaoshui Huang , Liang Xiao

Understanding user intent is essential for situational and context-aware decision-making. Motivated by a real-world scenario, this work addresses intent predictions of smart device users in the vicinity of vehicles by modeling sequential…

Editing objects within a scene is a critical functionality required across a broad spectrum of applications in computer vision and graphics. As 3D Gaussian Splatting (3DGS) emerges as a frontier in scene representation, the effective…

计算机视觉与模式识别 · 计算机科学 2024-06-04 Teng Xu , Jiamin Chen , Peng Chen , Youjia Zhang , Junqing Yu , Wei Yang

Accurate text recognition in low-light environments is essential for intelligent systems in applications ranging from autonomous vehicles to smart surveillance. However, challenges such as poor illumination and noise interference remain…

计算机视觉与模式识别 · 计算机科学 2026-04-28 Xuanshuo Fu , Lei Kang , Ernest Valveny , Dimosthenis Karatzas , Javier Vazquez-Corral

Scene text retrieval has made significant progress with the assistance of accurate text localization. However, existing approaches typically require costly bounding box annotations for training. Besides, they mostly adopt a customized…

计算机视觉与模式识别 · 计算机科学 2025-12-23 Liang Yin , Xudong Xie , Zhang Li , Xiang Bai , Yuliang Liu

Recent advances in Vision-Language-Action (VLA) models, powered by large language models and reinforcement learning-based fine-tuning, have shown remarkable progress in robotic manipulation. Existing methods often treat long-horizon actions…

机器人学 · 计算机科学 2025-12-25 Feng Xu , Guangyao Zhai , Xin Kong , Tingzhong Fu , Daniel F. N. Gordon , Xueli An , Benjamin Busam

Semantic Textual Similarity (STS) is a crucial component of many Natural Language Processing (NLP) applications. However, existing approaches typically reduce semantic nuances to a single score, limiting interpretability. To address this,…

计算与语言 · 计算机科学 2026-05-15 Diego Miguel Lozano , Daryna Dementieva , Alexander Fraser

The diversity in length constitutes a significant characteristic of text. Due to the long-tail distribution of text lengths, most existing methods for scene text recognition (STR) only work well on short or seen-length text, lacking the…

计算机视觉与模式识别 · 计算机科学 2023-08-25 Changxu Cheng , Peng Wang , Cheng Da , Qi Zheng , Cong Yao

We introduce DiffSteISR, a pioneering framework for reconstructing real-world stereo images. DiffSteISR utilizes the powerful prior knowledge embedded in pre-trained text-to-image model to efficiently recover the lost texture details in…

计算机视觉与模式识别 · 计算机科学 2024-08-16 Yuanbo Zhou , Xinlin Zhang , Wei Deng , Tao Wang , Tao Tan , Qinquan Gao , Tong Tong

Semantic Scene Completion (SSC) aims to perform geometric completion and semantic segmentation simultaneously. Despite the promising results achieved by existing studies, the inherently ill-posed nature of the task presents significant…

计算机视觉与模式识别 · 计算机科学 2024-11-21 Hyun-Kurl Jang , Jihun Kim , Hyeokjun Kweon , Kuk-Jin Yoon

Text recognition is an inherent integration of vision and language, encompassing the visual texture in stroke patterns and the semantic context among the character sequences. Towards advanced text recognition, there are three key…

计算机视觉与模式识别 · 计算机科学 2024-09-19 Humen Zhong , Zhibo Yang , Zhaohai Li , Peng Wang , Jun Tang , Wenqing Cheng , Cong Yao

Due to the high potential for abuse of GenAI systems, the task of detecting synthetic images has recently become of great interest to the research community. Unfortunately, existing image-space detectors quickly become obsolete as new…

计算机视觉与模式识别 · 计算机科学 2024-06-14 George Cazenavette , Avneesh Sud , Thomas Leung , Ben Usman

Scale variation is a deep-rooted problem in object counting, which has not been effectively addressed by existing scale-aware algorithms. An important factor is that they typically involve cooperative learning across multi-resolutions,…

计算机视觉与模式识别 · 计算机科学 2023-08-22 Tao Han , Lei Bai , Lingbo Liu , Wanli Ouyang
‹ 上一页 1 8 9 10 下一页 ›