中文
相关论文

相关论文: On-Device Document Classification using multimodal…

200 篇论文

We present TextMonkey, a large multimodal model (LMM) tailored for text-centric tasks. Our approach introduces enhancement across several dimensions: By adopting Shifted Window Attention with zero-initialization, we achieve cross-window…

计算机视觉与模式识别 · 计算机科学 2024-03-18 Yuliang Liu , Biao Yang , Qiang Liu , Zhang Li , Zhiyin Ma , Shuo Zhang , Xiang Bai

This paper highlights the need to bring document classification benchmarking closer to real-world applications, both in the nature of data tested ($X$: multi-channel, multi-paged, multi-industry; $Y$: class distributions and label set…

计算机视觉与模式识别 · 计算机科学 2023-11-01 Jordy Van Landeghem , Sanket Biswas , Matthew B. Blaschko , Marie-Francine Moens

In spite of the high accuracy of the existing optical mark reading (OMR) systems and devices, a few restrictions remain existent. In this work, we aim to reduce the restrictions of multiple choice questions (MCQ) within tests. We use an…

计算机视觉与模式识别 · 计算机科学 2019-01-15 Mahmoud Afifi , Khaled F. Hussain

Deploying deep learning (DL) on mobile devices has been a notable trend in recent years. To support fast inference of on-device DL, DL libraries play a critical role as algorithms and hardware do. Unfortunately, no prior work ever dives…

机器学习 · 计算机科学 2022-07-07 Qiyang Zhang , Xiang Li , Xiangying Che , Xiao Ma , Ao Zhou , Mengwei Xu , Shangguang Wang , Yun Ma , Xuanzhe Liu

Digitization of medical records often relies on smartphone photographs of printed reports, producing images degraded by blur, shadows, and other noise. Conventional OCR systems, optimized for clean scans, perform poorly under such…

信息检索 · 计算机科学 2025-11-18 Nikita Neveditsin , Pawan Lingras , Salil Patil , Swarup Patil , Vijay Mago

Mobile smartphones along with embedded sensors have become an efficient enabler for various mobile applications including opportunistic sensing. The hi-tech advances in smartphones are opening up a world of possibilities. This paper…

网络与互联网体系结构 · 计算机科学 2014-05-23 Prem Prakash Jayaraman , Charith Perera , Dimitrios Georgakopoulos , Arkady Zaslavsky

Document classification forms the backbone of modern enterprise content management, yet existing benchmarks remain trapped in oversimplified paradigms -- single domain settings with flat label structures -- that bear little resemblance to…

计算与语言 · 计算机科学 2026-05-15 Denghao Ma , Qing Liu , Zulong Chen , Chuanfei Xu , Jia Xu , Zhibo Yang , Wei Shao , Zhao Li

Large Language Model (LLM) pre-training exhausts an ever growing compute budget, yet recent research has demonstrated that careful document selection enables comparable model quality with only a fraction of the FLOPs. Inspired by efforts…

计算与语言 · 计算机科学 2024-06-10 Xiang Kong , Tom Gunter , Ruoming Pang

Optical Character Recognition (OCR) for data extraction from documents is essential to intelligent informatics, such as digitizing medical records and recognizing road signs. Multi-modal Large Language Models (LLMs) can solve this task and…

计算机视觉与模式识别 · 计算机科学 2026-01-06 Hyakka Nakada , Yoshiyasu Tanaka

Classifying journals or publications into research areas is an essential element of many bibliometric analyses. Classification usually takes place at the level of journals, where the Web of Science subject categories are the most popular…

数字图书馆 · 计算机科学 2012-03-05 Ludo Waltman , Nees Jan van Eck

Despite the growing adoption of electronic health records, many processes still rely on paper documents, reflecting the heterogeneous real-world conditions in which healthcare is delivered. The manual transcription process is time-consuming…

There are a plethora of methods and algorithms that solve the classical multi-label document classification. However, when it comes to deployment and usage in an industry setting, most, if not all the contemporary approaches fail to address…

计算与语言 · 计算机科学 2023-01-18 Arshad Javeed

Determining the sentence similarity between Short Message Service (SMS) texts/sentences plays a significant role in mobile device industry. Gauging the similarity between SMS data is thus necessary for various applications like enhanced…

计算与语言 · 计算机科学 2022-01-03 Arun D Prabhu , Nikhil Arora , Shubham Vatsal , Gopi Ramena , Sukumar Moharana , Naresh Purre

Sharing location traces with context-aware service providers has privacy implications. Location-privacy preserving mechanisms, such as obfuscation, anonymization and cryptographic primitives, have been shown to have impractical…

密码学与安全 · 计算机科学 2018-02-21 Vaibhav Kulkarni , Arielle Moro , Bertil Chapuis , Benoit Garbinato

Meta-learning has emerged as a prominent technology for few-shot text classification and has achieved promising performance. However, existing methods often encounter difficulties in drawing accurate class prototypes from support set…

计算机视觉与模式识别 · 计算机科学 2024-10-15 Xinyue Liu , Yunlong Gao , Linlin Zong , Bo Xu

The presence of any type of defect on the glass screen of smart devices has a great impact on their quality. We present a robust semi-supervised learning framework for intelligent micro-scaled localization and classification of defects on a…

计算机视觉与模式识别 · 计算机科学 2020-10-05 M Usman Maqbool Bhutta , Shoaib Aslam , Peng Yun , Jianhao Jiao , Ming Liu

An automatic document classification system is presented that detects textual content in images and classifies documents into four predefined categories (Invoice, Report, Letter, and Form). The system supports both offline images (e.g.,…

计算机视觉与模式识别 · 计算机科学 2025-10-16 Aya Kaysan Bahjat

Recommender systems have been widely deployed in various real-world applications to help users identify content of interest from massive amounts of information. Traditional recommender systems work by collecting user-item interaction data…

信息检索 · 计算机科学 2025-08-07 Hongzhi Yin , Liang Qu , Tong Chen , Wei Yuan , Ruiqi Zheng , Jing Long , Xin Xia , Yuhui Shi , Chengqi Zhang

Diacritic characters can be considered as a unique set of characters providing us with adequate and significant clue in identifying a given language with considerably high accuracy. Diacritics, though associated with phonetics often serve…

计算机视觉与模式识别 · 计算机科学 2021-12-07 Shubham Vatsal , Nikhil Arora , Gopi Ramena , Sukumar Moharana , Dhruval Jain , Naresh Purre , Rachit S Munjal

We propose a novel end-to-end solution that performs a Hierarchical Layout Analysis of screenshots and document images on resource constrained devices like mobilephones. Our approach segments entities like Grid, Image, Text and Icon blocks…

计算机视觉与模式识别 · 计算机科学 2021-04-22 Manoj Goyal , Rachit S Munjal , Sukumar Moharana , Deepak Garg , Debi Prasanna Mohanty , Siva Prasad Thota