中文
相关论文

相关论文: On-Device Document Classification using multimodal…

200 篇论文

In this paper, we address the problem of classifying documents available from the global network of (open access) repositories according to their type. We show that the metadata provided by repositories enabling us to distinguish research…

数字图书馆 · 计算机科学 2017-07-14 Aristotelis Charalampous , Petr Knoth

This demo presents a novel end-to-end framework that combines on-device large language models (LLMs) with smartphone sensing technologies to achieve context-aware and personalized services. The framework addresses critical limitations of…

人机交互 · 计算机科学 2024-07-25 Shiquan Zhang , Ying Ma , Le Fang , Hong Jia , Simon D'Alfonso , Vassilis Kostakos

Knowledge of source smartphone corresponding to a document image can be helpful in a variety of applications including copyright infringement, ownership attribution, leak identification and usage restriction. In this letter, we investigate…

多媒体 · 计算机科学 2019-06-18 Sharad Joshi , Suraj Saxena , Nitin Khanna

Academic research tends to focus on new models for document understanding creating a wide gap in the literature between model definition and running models at production scale. To close that gap, we present a microservice architecture that…

While small language models (SLMs) show promises for mobile deployment, their real-world performance and applications on smartphones remains underexplored. We present SlimLM, a series of SLMs optimized for document assistance tasks on…

计算与语言 · 计算机科学 2024-11-27 Thang M. Pham , Phat T. Nguyen , Seunghyun Yoon , Viet Dac Lai , Franck Dernoncourt , Trung Bui

With the rapid development of storage and computing power on mobile devices, it becomes critical and popular to deploy models on devices to save onerous communication latencies and to capture real-time features. While quite a lot of works…

机器学习 · 计算机科学 2021-06-18 Jiangchao Yao , Feng Wang , KunYang Jia , Bo Han , Jingren Zhou , Hongxia Yang

A growing number of commercially available mobile phones come with integrated high-resolution digital cameras. That enables a new class of dedicated applications to image analysis such as mobile visual search, image cropping, object…

计算机视觉与模式识别 · 计算机科学 2021-08-30 Alessandro Bruno

Smartphone clip-on microscopes turn everyday devices into low-cost, portable imaging systems that can even reveal fungal structures at the microscopic level, enabling mold inspection beyond unaided visual checks. In this paper, we introduce…

计算机视觉与模式识别 · 计算机科学 2026-04-13 Dinh Nam Pham , Leonard Prokisch , Bennet Meyer , Jonas Thumbs

Document content extraction is a critical task in computer vision, underpinning the data needs of large language models (LLMs) and retrieval-augmented generation (RAG) systems. Despite recent progress, current document parsing methods have…

While the incipient internet was largely text-based, the modern digital world is becoming increasingly multi-modal. Here, we examine multi-modal classification where one modality is discrete, e.g. text, and the other is continuous, e.g.…

计算与语言 · 计算机科学 2018-02-09 D. Kiela , E. Grave , A. Joulin , T. Mikolov

Classification of document images is a critical step for archival of old manuscripts, online subscription and administrative procedures. Computer vision and deep learning have been suggested as a first solution to classify documents based…

计算机视觉与模式识别 · 计算机科学 2019-07-16 Nicolas Audebert , Catherine Herold , Kuider Slimani , Cédric Vidal

We introduce the Brno Mobile OCR Dataset (B-MOD) for document Optical Character Recognition from low-quality images captured by handheld mobile devices. While OCR of high-quality scanned documents is a mature field where many commercial…

计算机视觉与模式识别 · 计算机科学 2019-07-03 Martin Kišš , Michal Hradiš , Oldřich Kodym

In modern mobile applications, users frequently encounter various new contexts, necessitating on-device continual learning (CL) to ensure consistent model performance. While existing research predominantly focused on developing lightweight…

机器学习 · 计算机科学 2024-10-25 Chen Gong , Zhenzhe Zheng , Fan Wu , Xiaofeng Jia , Guihai Chen

Large Multimodal Models (LMMs) have recently shown strong performance on Optical Character Recognition (OCR) tasks, demonstrating their promising capability in document literacy. However, their effectiveness in real-world applications…

On-device recommendation is critical for a number of real-world applications, especially in scenarios that have agreements on execution latency, user privacy, and robust functionality when internet connectivity is unstable or even…

信息检索 · 计算机科学 2026-01-15 Xin Xia , Hongzhi Yin , Shane Culpepper

Document understanding tasks, in particular, Visually-rich Document Entity Retrieval (VDER), have gained significant attention in recent years thanks to their broad applications in enterprise AI. However, publicly available data have been…

计算与语言 · 计算机科学 2023-10-27 Lijun Yu , Jin Miao , Xiaoyu Sun , Jiayi Chen , Alexander G. Hauptmann , Hanjun Dai , Wei Wei

On-device machine learning (ML) promises to improve the privacy, responsiveness, and proliferation of new, intelligent user experiences by moving ML computation onto everyday personal devices. However, today's large ML models must be…

人机交互 · 计算机科学 2024-04-05 Fred Hohman , Mary Beth Kery , Donghao Ren , Dominik Moritz

Over the past decade, machine learning methods have given us driverless cars, voice recognition, effective web search, and a much better understanding of the human genome. Machine learning is so common today that it is used dozens of times…

计算机视觉与模式识别 · 计算机科学 2021-06-22 Omer Aydin

Page classification is a crucial component to any document analysis system, allowing for complex branching control flows for different components of a given document. Utilizing both the visual and textual content of a page, the proposed…

计算机视觉与模式识别 · 计算机科学 2019-12-11 Tyler Dauphinee , Nikunj Patel , Mohammad Rashidi

The demand for on-device document recognition systems increases in conjunction with the emergence of more strict privacy and security requirements. In such systems, there is no data transfer from the end device to a third-party information…

计算机视觉与模式识别 · 计算机科学 2021-09-27 D. V. Tropin , A. M. Ershov , D. P. Nikolaev , V. V. Arlazarov