中文
相关论文

相关论文: ABot-OCR Technical Report

200 篇论文

DeepSeek-OCR leverages visual-text compression to reduce long-text processing costs and accelerate inference, yet visual tokens remain prone to redundant textual and structural information. Moreover, current token pruning methods for…

计算机视觉与模式识别 · 计算机科学 2026-05-22 Ben Wan , Yan Feng , Zihan Tang , Weizhe Huang , Yuting Zeng , Jia Wang , Tongxuan Liu

The automatic recognition of tabular data in document images presents a significant challenge due to the diverse range of table styles and complex structures. Tables offer valuable content representation, enhancing the predictive…

计算机视觉与模式识别 · 计算机科学 2024-04-22 Avinash Anand , Raj Jaiswal , Pijush Bhuyan , Mohit Gupta , Siddhesh Bangar , Md. Modassir Imam , Rajiv Ratn Shah , Shin'ichi Satoh

We introduce OneCAT, a unified multimodal model that seamlessly integrates understanding, generation, and editing within a novel, pure decoder-only transformer architecture. Our framework uniquely eliminates the need for external components…

计算机视觉与模式识别 · 计算机科学 2025-10-08 Han Li , Xinyu Peng , Yaoming Wang , Zelin Peng , Xin Chen , Rongxiang Weng , Jingang Wang , Xunliang Cai , Wenrui Dai , Hongkai Xiong

Document parsing is a fine-grained task where image resolution significantly impacts performance. While advanced research leveraging vision-language models benefits from high-resolution input to boost model performance, this often leads to…

Recent advances in OCR have shown that an end-to-end (E2E) training pipeline that includes both detection and recognition leads to the best results. However, many existing methods focus primarily on Latin-alphabet languages, often even only…

计算机视觉与模式识别 · 计算机科学 2021-03-31 Jing Huang , Guan Pang , Rama Kovvuri , Mandy Toh , Kevin J Liang , Praveen Krishnan , Xi Yin , Tal Hassner

Acoustic-to-Word recognition provides a straightforward solution to end-to-end speech recognition without needing external decoding, language model re-scoring or lexicon. While character-based models offer a natural solution to the…

音频与语音处理 · 电气工程与系统科学 2018-08-22 Shruti Palaskar , Florian Metze

This paper presents an end-to-end suite for multilingual information extraction and processing from image-based documents. The system uses Optical Character Recognition (Tesseract) to extract text in languages such as English, Hindi, and…

计算与语言 · 计算机科学 2025-05-19 Hrishit Madhavi , Jacob Cherian , Yuvraj Khamkar , Dhananjay Bhagat

Optical Coherence Tomography Angiography (OCTA) and its derived en-face projections provide high-resolution visualization of the retinal and choroidal vasculature, which is critical for the rapid and accurate diagnosis of retinal diseases.…

计算机视觉与模式识别 · 计算机科学 2025-09-10 Pooya Khosravi , Kun Han , Anthony T. Wu , Arghavan Rezvani , Zexin Feng , Xiaohui Xie

Optical Coherence Tomography Angiography (OCTA) is a non-invasive and non-contacting imaging technique providing visualization of microvasculature of retina and optic nerve head in human eyes in vivo. The adequate image quality of OCTA is…

图像与视频处理 · 电气工程与系统科学 2021-07-23 Yufei Wang , Yiqing Shen , Meng Yuan , Jing Xu , Bin Yang , Chi Liu , Wenjia Cai , Weijing Cheng , Wei Wang

Recent advances in representation learning often rely on holistic embeddings that entangle multiple semantic components, limiting interpretability and generalization. These issues are especially critical in medical imaging, where downstream…

计算机视觉与模式识别 · 计算机科学 2025-11-21 Sifan Song , Siyeop Yoon , Pengfei Jin , Sekeun Kim , Matthew Tivnan , Yujin Oh , Runqi Meng , Ling Chen , Zhiliang Lyu , Dufan Wu , Ning Guo , Xiang Li , Quanzheng Li

With the rapid advancement of digitalization, various document images are being applied more extensively in production and daily life, and there is an increasingly urgent need for fast and accurate parsing of the content in document images.…

计算机视觉与模式识别 · 计算机科学 2025-06-27 Feng Ni , Kui Huang , Yao Lu , Wenyu Lv , Guanzhong Wang , Zeyu Chen , Yi Liu

A great deal of historical corpora suffer from errors introduced by the OCR (optical character recognition) methods used in the digitization process. Correcting these errors manually is a time-consuming process and a great part of the…

计算与语言 · 计算机科学 2020-07-23 Mika Hämäläinen , Simon Hengchen

Controllable captioning is essential for precise multimodal alignment and instruction following, yet existing models often lack fine-grained control and reliable evaluation protocols. To address this gap, we present the AnyCap Project, an…

计算机视觉与模式识别 · 计算机科学 2025-10-29 Yiming Ren , Zhiqiang Lin , Yu Li , Gao Meng , Weiyun Wang , Junjie Wang , Zicheng Lin , Jifeng Dai , Yujiu Yang , Wenhai Wang , Ruihang Chu

Recent advances in large language models (LLMs) enable agentic systems trained with reinforcement learning (RL) over multi-turn interaction trajectories, but practical deployment is bottlenecked by rapidly growing textual histories that…

机器学习 · 计算机科学 2026-03-03 Lang Feng , Fuchao Yang , Feng Chen , Xin Cheng , Haiyang Xu , Zhenglin Wan , Ming Yan , Bo An

Optical Coherence Tomography (OCT) is a widely used non-invasive biomedical imaging modality that can rapidly provide volumetric images of samples. Here, we present a deep learning-based image reconstruction framework that can generate…

图像与视频处理 · 电气工程与系统科学 2021-07-30 Yijie Zhang , Tairan Liu , Manmohan Singh , Yilin Luo , Yair Rivenson , Kirill V. Larin , Aydogan Ozcan

Good OCR results for historical printings rely on the availability of recognition models trained on diplomatic transcriptions as ground truth, which is both a scarce resource and time-consuming to generate. Instead of having to train a…

数字图书馆 · 计算机科学 2016-10-21 U. Springmann , F. Fink , K. U. Schulz

Modern LVLMs still struggle to achieve fine-grained document understanding, such as OCR/translation/caption for regions of interest to the user, tasks that require the context of the entire page, or even multiple pages. Accordingly, this…

计算机视觉与模式识别 · 计算机科学 2024-05-24 Chenglong Liu , Haoran Wei , Jinyue Chen , Lingyu Kong , Zheng Ge , Zining Zhu , Liang Zhao , Jianjian Sun , Chunrui Han , Xiangyu Zhang

Documents are central to many business systems, and include forms, reports, contracts, invoices or purchase orders. The information in documents is typically in natural language, but can be organized in various layouts and formats. There…

信息检索 · 计算机科学 2021-10-08 Sumit Shekhar , Bhanu Prakash Reddy Guda , Ashutosh Chaubey , Ishan Jindal , Avneet Jain

End-to-end autonomous driving has great potential in the transportation industry. However, the lack of transparency and interpretability of the automatic decision-making process hinders its industrial adoption in practice. There have been…

计算机视觉与模式识别 · 计算机科学 2023-02-02 Bu Jin , Xinyu Liu , Yupeng Zheng , Pengfei Li , Hao Zhao , Tong Zhang , Yuhang Zheng , Guyue Zhou , Jingjing Liu

AEC drawings encode geometry and semantics through symbols, layout conventions, and dense annotation, yet it remains unclear whether modern multimodal and vision-language models can reliably interpret this graphical language. We present…

人工智能 · 计算机科学 2026-01-09 Aleksei Kondratenko , Mussie Birhane , Houssame E. Hsain , Guido Maciocci