中文
相关论文

相关论文: Ensemble of Anchor-Free Models for Robust Bangla D…

200 篇论文

The current digital environment is characterized by the widespread presence of data, particularly unstructured data, which poses many issues in sectors including finance, healthcare, and education. Conventional techniques for data…

计算机视觉与模式识别 · 计算机科学 2023-10-02 Herman Sugiharto , Yorissa Silviana , Yani Siti Nurpazrin

Segmentation of handwritten document images into text lines and words is one of the most significant and challenging tasks in the development of a complete Optical Character Recognition (OCR) system. This paper addresses the automatic…

计算机视觉与模式识别 · 计算机科学 2020-09-18 Pawan Kumar Singh , Shubham Sinha , Sagnik Pal Chowdhury , Ram Sarkar , Mita Nasipuri

Document parsing has garnered widespread attention as vision-language models (VLMs) advance OCR capabilities. However, the field remains fragmented across dozens of specialized models with varying strengths, forcing users to navigate…

计算机视觉与模式识别 · 计算机科学 2026-02-06 Hao Feng , Wei Shi , Ke Zhang , Xiang Fei , Lei Liao , Dingkang Yang , Yongkun Du , Xuecheng Wu , Jingqun Tang , Yang Liu , Hong Chen , Can Huang

While strides have been made in deep learning based Bengali Optical Character Recognition (OCR) in the past decade, the absence of large Document Layout Analysis (DLA) datasets has hindered the application of OCR in document transcription,…

This paper presents a high-quality dataset for evaluating the quality of Bangla word embeddings, which is a fundamental task in the field of Natural Language Processing (NLP). Despite being the 7th most-spoken language in the world, Bangla…

计算与语言 · 计算机科学 2023-04-11 Mousumi Akter , Souvika Sarkar , Shubhra Kanti Karmaker Santu

Accurate layout analysis without subsequent text-line segmentation remains an ongoing challenge, especially when facing the Kangyur, a kind of historical Tibetan document featuring considerable touching components and mottled background.…

计算机视觉与模式识别 · 计算机科学 2024-10-28 Penghai Zhao , Weilan Wang , Zhengqi Cai , Guowei Zhang , Yuqi Lu

The reliable identification of mitotic figures in whole-slide histopathological images remains difficult, owing to their low prevalence, substantial morphological heterogeneity, and the inconsistencies introduced by tissue processing and…

图像与视频处理 · 电气工程与系统科学 2025-09-23 Navya Sri Kelam , Akash Parekh , Saikiran Bonthu , Nitin Singhal

As obesity rates continue to increase, automated calorie tracking has become a vital tool for people seeking to maintain a healthy lifestyle or adhere to a diet plan. Although numerous research efforts have addressed this issue, existing…

计算机视觉与模式识别 · 计算机科学 2025-09-03 Aparup Dhar , MD Tamim Hossain , Pritom Barua

Text-rich document understanding (TDU) requires comprehensive analysis of documents containing substantial textual content and complex layouts. While Multimodal Large Language Models (MLLMs) have achieved fast progress in this domain,…

计算机视觉与模式识别 · 计算机科学 2025-03-20 Wenhui Liao , Jiapeng Wang , Hongliang Li , Chengyu Wang , Jun Huang , Lianwen Jin

In order to ensure traffic safety through a reduction in fatalities and accidents, vehicle speed detection is essential. Relentless driving practices are discouraged by the enforcement of speed restrictions, which are made possible by…

计算机视觉与模式识别 · 计算机科学 2024-08-31 SM Shaqib , Alaya Parvin Alo , Shahriar Sultan Ramit , Afraz Ul Haque Rupak , Sadman Sadik Khan , Md. Sadekur Rahman

This study explores a comprehensive approach to obstacle detection using advanced YOLO models, specifically YOLOv8, YOLOv7, YOLOv6, and YOLOv5. Leveraging deep learning techniques, the research focuses on the performance comparison of these…

计算机视觉与模式识别 · 计算机科学 2024-10-15 Santiago Pérez , Camila Gómez , Matías Rodríguez

Large language models work well for technical problem solving in English but perform poorly when the same questions are asked in Bangla. A simple solution would be to translate Bangla questions into English first and then use these models.…

计算与语言 · 计算机科学 2025-11-06 Kazi Reyazul Hasan , Mubasshira Musarrat , A. B. M. Alim Al Islam , Muhammad Abdullah Adnan

As urbanization speeds up and traffic flow increases, the issue of pavement distress is becoming increasingly pronounced, posing a severe threat to road safety and service life. Traditional methods of pothole detection rely on manual…

计算机视觉与模式识别 · 计算机科学 2025-07-17 Haomin Zuo , Zhengyang Li , Jiangchuan Gong , Zhen Tian

Rapid urbanization in megacities around the world, like Dhaka, has caused numerous transportation challenges that need to be addressed. Emerging technologies of deep learning and artificial intelligence can help us solve these problems to…

计算机视觉与模式识别 · 计算机科学 2024-10-15 Shahriar Ahmad Fahim

The task of locating and classifying different types of vehicles has become a vital element in numerous applications of automation and intelligent systems ranging from traffic surveillance to vehicle identification and many more. In recent…

The study involves a comprehensive performance analysis of popular classification and segmentation models, applied over a Bangladeshi pothole dataset, being developed by the authors of this research. This custom dataset of 824 samples,…

Today all kind of information is getting digitized and along with all this digitization, the huge archive of various kinds of documents is being digitized too. We know that, Optical Character Recognition is the method through which,…

计算机视觉与模式识别 · 计算机科学 2017-01-31 Md. Fahad Hasan , Tasmin Afroz , Sabir Ismail , Md. Saiful Islam

This paper presents MSLEF, a multi-segment ensemble framework that employs LLM fine-tuning to enhance resume parsing in recruitment automation. It integrates fine-tuned Large Language Models (LLMs) using weighted voting, with each model…

计算与语言 · 计算机科学 2025-09-09 Omar Walid , Mohamed T. Younes , Khaled Shaban , Mai Hassan , Ali Hamdi

This paper presents a method for detecting grammatical errors in Bangla using a Text-to-Text Transfer Transformer (T5) Language Model, using the small variant of BanglaT5, fine-tuned on a corpus of 9385 sentences where errors were bracketed…

计算与语言 · 计算机科学 2023-03-21 H. A. Z. Sameen Shahgir , Khondker Salman Sayeed

This paper introduces a comprehensive approach for segmenting regions of interest (ROI) in diverse medical imaging datasets, encompassing ultrasound, CT scans, and X-ray images. The proposed method harnesses the capabilities of the YOLOv8…

计算机视觉与模式识别 · 计算机科学 2023-10-23 Sumit Pandey , Kuan-Fu Chen , Erik B. Dam