中文
相关论文

相关论文: Authorship Attribution in Bangla literature using …

200 篇论文

Authorship Attribution is a long-standing problem in Natural Language Processing. Several statistical and computational methods have been used to find a solution to this problem. In this paper, we have proposed methods to deal with the…

计算与语言 · 计算机科学 2016-09-08 Shanta Phani , Shibamouli Lahiri , Arindam Biswas

The Bengali language is the 5th most spoken native and 7th most spoken language in the world, and Bengali handwritten character recognition has attracted researchers for decades. However, other languages such as English, Arabic, Turkey, and…

计算机视觉与模式识别 · 计算机科学 2024-08-21 Farhanul Haque , Md. Al-Hasan , Sumaiya Tabssum Mou , Abu Saleh Musa Miah , Jungpil Shin , Md Abdur Rahim

Classification techniques for images of handwritten characters are susceptible to noise. Quadtrees can be an efficient representation for learning from sparse features. In this paper, we improve the effectiveness of probabilistic quadtrees…

计算机视觉与模式识别 · 计算机科学 2018-06-22 Manohar Karki , Qun Liu , Robert DiBiano , Saikat Basu , Supratik Mukhopadhyay

Sequence labeling architectures use word embeddings for capturing similarity, but suffer when handling previously unseen or rare words. We investigate character-level extensions to such models and propose a novel architecture for combining…

计算与语言 · 计算机科学 2016-11-15 Marek Rei , Gamal K. O. Crichton , Sampo Pyysalo

Authorship attribution is the process of identifying the author of a text. Approaches to tackling it have been conventionally divided into classification-based ones, which work well for small numbers of candidate authors, and…

计算与语言 · 计算机科学 2021-05-18 Chakaveh Saedi , Mark Dras

Identification of minimum number of local regions of a handwritten character image, containing well-defined discriminating features which are sufficient for a minimal but complete description of the character is a challenging task. A new…

计算机视觉与模式识别 · 计算机科学 2016-05-03 Ritesh Sarkhel , Amit K Saha , Nibaran Das

Large language models have shown breakthrough potential in many NLP domains. Here we consider their use for stylometry, specifically authorship identification in Early Modern English drama. We find both promising and concerning results;…

计算与语言 · 计算机科学 2023-10-31 Rebecca M. M. Hicke , David Mimno

Syllabification does not seem to improve word-level RNN language modeling quality when compared to character-based segmentation. However, our best syllable-aware language model, achieving performance comparable to the competitive…

计算与语言 · 计算机科学 2017-07-21 Zhenisbek Assylbekov , Rustem Takhanov , Bagdat Myrzakhmetov , Jonathan N. Washington

Authorship attribution models fine-tuned with the same pretrained encoder, data, and loss can differ four-fold in performance depending only on their scoring mechanism. We use mechanistic interpretability tools to explain this gap.…

计算与语言 · 计算机科学 2026-05-27 Francis Kulumba , Guillaume Vimont , Laurent Romary , Florian Cafiero

The writing style of a person can be affirmed as a unique identity indicator; the words used, and the structuring of the sentences are clear measures which can identify the author of a specific work. Stylometry and its subset - Authorship…

计算与语言 · 计算机科学 2018-12-27 Abhay Sharma , Ananya Nandan , Reetika Ralhan

Authorship analysis has traditionally focused on lexical and stylistic cues within text, while higher-level narrative structure remains underexplored, particularly for low-resource languages such as Urdu. This work proposes a graph-based…

计算与语言 · 计算机科学 2025-12-16 Hassan Mujtaba , Hamza Naveed , Hanzlah Munir

In this work, a novel deep learning technique for the recognition of handwritten Bangla isolated compound character is presented and a new benchmark of recognition accuracy on the CMATERdb 3.1.3.3 dataset is reported. Greedy layer wise…

计算机视觉与模式识别 · 计算机科学 2018-02-05 Saikat Roy , Nibaran Das , Mahantapas Kundu , Mita Nasipuri

Word-level handwritten optical character recognition (OCR) remains a challenge for morphologically rich languages like Bangla. The complexity arises from the existence of a large number of alphabets, the presence of several diacritic forms,…

计算机视觉与模式识别 · 计算机科学 2022-05-24 Md. Ismail Hossain , Mohammed Rakib , Sabbir Mollah , Fuad Rahman , Nabeel Mohammed

Text normalization is an important enabling technology for several NLP tasks. Recently, neural-network-based approaches have outperformed well-established models in this task. However, in languages other than English, there has been little…

计算与语言 · 计算机科学 2018-09-06 Daniel Watson , Nasser Zalmout , Nizar Habash

Finding local invariant patterns in handwrit-ten characters and/or digits for optical character recognition is a difficult task. Variations in writing styles from one person to another make this task challenging. We have proposed a…

计算机视觉与模式识别 · 计算机科学 2020-04-28 Animesh Singh , Ritesh Sarkhel , Nibaran Das , Mahantapas Kundu , Mita Nasipuri

This study investigates the performance of few-shot learning (FSL) approaches in recognizing Bangla handwritten characters and numerals using limited labeled data. It demonstrates the applicability of these methods to scripts with intricate…

计算机视觉与模式识别 · 计算机科学 2025-06-03 Mehedi Ahamed , Radib Bin Kabir , Tawsif Tashwar Dipto , Mueeze Al Mushabbir , Sabbir Ahmed , Md. Hasanul Kabir

The study of dreams has been central to understanding human (un)consciousness, cognition, and culture for centuries. Analyzing dreams quantitatively depends on labor-intensive, manual annotation of dream narratives. We automate this process…

计算与语言 · 计算机科学 2024-03-26 Gustave Cortal

Due to digitalization in everyday life, the need for automatically recognizing handwritten digits is increasing. Handwritten digit recognition is essential for numerous applications in various industries. Bengali ranks the fifth largest…

Subword-level models have been the dominant paradigm in NLP. However, character-level models have the benefit of seeing each character individually, providing the model with more detailed information that ultimately could lead to better…

计算与语言 · 计算机科学 2022-12-05 Lukas Edman , Antonio Toral , Gertjan van Noord

Current models for quotation attribution in literary novels assume varying levels of available information in their training and test data, which poses a challenge for in-the-wild inference. Here, we approach quotation attribution as a set…

计算与语言 · 计算机科学 2023-07-10 Krishnapriya Vishnubhotla , Frank Rudzicz , Graeme Hirst , Adam Hammond