English
Related papers

Related papers: Authorship Attribution in Bangla literature using …

200 papers

Authorship Attribution is a long-standing problem in Natural Language Processing. Several statistical and computational methods have been used to find a solution to this problem. In this paper, we have proposed methods to deal with the…

Computation and Language · Computer Science 2016-09-08 Shanta Phani , Shibamouli Lahiri , Arindam Biswas

The Bengali language is the 5th most spoken native and 7th most spoken language in the world, and Bengali handwritten character recognition has attracted researchers for decades. However, other languages such as English, Arabic, Turkey, and…

Computer Vision and Pattern Recognition · Computer Science 2024-08-21 Farhanul Haque , Md. Al-Hasan , Sumaiya Tabssum Mou , Abu Saleh Musa Miah , Jungpil Shin , Md Abdur Rahim

Classification techniques for images of handwritten characters are susceptible to noise. Quadtrees can be an efficient representation for learning from sparse features. In this paper, we improve the effectiveness of probabilistic quadtrees…

Computer Vision and Pattern Recognition · Computer Science 2018-06-22 Manohar Karki , Qun Liu , Robert DiBiano , Saikat Basu , Supratik Mukhopadhyay

Sequence labeling architectures use word embeddings for capturing similarity, but suffer when handling previously unseen or rare words. We investigate character-level extensions to such models and propose a novel architecture for combining…

Computation and Language · Computer Science 2016-11-15 Marek Rei , Gamal K. O. Crichton , Sampo Pyysalo

Authorship attribution is the process of identifying the author of a text. Approaches to tackling it have been conventionally divided into classification-based ones, which work well for small numbers of candidate authors, and…

Computation and Language · Computer Science 2021-05-18 Chakaveh Saedi , Mark Dras

Identification of minimum number of local regions of a handwritten character image, containing well-defined discriminating features which are sufficient for a minimal but complete description of the character is a challenging task. A new…

Computer Vision and Pattern Recognition · Computer Science 2016-05-03 Ritesh Sarkhel , Amit K Saha , Nibaran Das

Large language models have shown breakthrough potential in many NLP domains. Here we consider their use for stylometry, specifically authorship identification in Early Modern English drama. We find both promising and concerning results;…

Computation and Language · Computer Science 2023-10-31 Rebecca M. M. Hicke , David Mimno

Syllabification does not seem to improve word-level RNN language modeling quality when compared to character-based segmentation. However, our best syllable-aware language model, achieving performance comparable to the competitive…

Computation and Language · Computer Science 2017-07-21 Zhenisbek Assylbekov , Rustem Takhanov , Bagdat Myrzakhmetov , Jonathan N. Washington

Authorship attribution models fine-tuned with the same pretrained encoder, data, and loss can differ four-fold in performance depending only on their scoring mechanism. We use mechanistic interpretability tools to explain this gap.…

Computation and Language · Computer Science 2026-05-27 Francis Kulumba , Guillaume Vimont , Laurent Romary , Florian Cafiero

The writing style of a person can be affirmed as a unique identity indicator; the words used, and the structuring of the sentences are clear measures which can identify the author of a specific work. Stylometry and its subset - Authorship…

Computation and Language · Computer Science 2018-12-27 Abhay Sharma , Ananya Nandan , Reetika Ralhan

Authorship analysis has traditionally focused on lexical and stylistic cues within text, while higher-level narrative structure remains underexplored, particularly for low-resource languages such as Urdu. This work proposes a graph-based…

Computation and Language · Computer Science 2025-12-16 Hassan Mujtaba , Hamza Naveed , Hanzlah Munir

In this work, a novel deep learning technique for the recognition of handwritten Bangla isolated compound character is presented and a new benchmark of recognition accuracy on the CMATERdb 3.1.3.3 dataset is reported. Greedy layer wise…

Computer Vision and Pattern Recognition · Computer Science 2018-02-05 Saikat Roy , Nibaran Das , Mahantapas Kundu , Mita Nasipuri

Word-level handwritten optical character recognition (OCR) remains a challenge for morphologically rich languages like Bangla. The complexity arises from the existence of a large number of alphabets, the presence of several diacritic forms,…

Computer Vision and Pattern Recognition · Computer Science 2022-05-24 Md. Ismail Hossain , Mohammed Rakib , Sabbir Mollah , Fuad Rahman , Nabeel Mohammed

Text normalization is an important enabling technology for several NLP tasks. Recently, neural-network-based approaches have outperformed well-established models in this task. However, in languages other than English, there has been little…

Computation and Language · Computer Science 2018-09-06 Daniel Watson , Nasser Zalmout , Nizar Habash

Finding local invariant patterns in handwrit-ten characters and/or digits for optical character recognition is a difficult task. Variations in writing styles from one person to another make this task challenging. We have proposed a…

Computer Vision and Pattern Recognition · Computer Science 2020-04-28 Animesh Singh , Ritesh Sarkhel , Nibaran Das , Mahantapas Kundu , Mita Nasipuri

This study investigates the performance of few-shot learning (FSL) approaches in recognizing Bangla handwritten characters and numerals using limited labeled data. It demonstrates the applicability of these methods to scripts with intricate…

Computer Vision and Pattern Recognition · Computer Science 2025-06-03 Mehedi Ahamed , Radib Bin Kabir , Tawsif Tashwar Dipto , Mueeze Al Mushabbir , Sabbir Ahmed , Md. Hasanul Kabir

The study of dreams has been central to understanding human (un)consciousness, cognition, and culture for centuries. Analyzing dreams quantitatively depends on labor-intensive, manual annotation of dream narratives. We automate this process…

Computation and Language · Computer Science 2024-03-26 Gustave Cortal

Due to digitalization in everyday life, the need for automatically recognizing handwritten digits is increasing. Handwritten digit recognition is essential for numerous applications in various industries. Bengali ranks the fifth largest…

Computer Vision and Pattern Recognition · Computer Science 2023-05-23 Muntarin Islam , Shabbir Ahmed Shuvo , Musarrat Saberin Nipun , Rejwan Bin Sulaiman , Jannatul Nayeem , Zubaer Haque , Md Mostak Shaikh , Md Sakib Ullah Sourav

Subword-level models have been the dominant paradigm in NLP. However, character-level models have the benefit of seeing each character individually, providing the model with more detailed information that ultimately could lead to better…

Computation and Language · Computer Science 2022-12-05 Lukas Edman , Antonio Toral , Gertjan van Noord

Current models for quotation attribution in literary novels assume varying levels of available information in their training and test data, which poses a challenge for in-the-wild inference. Here, we approach quotation attribution as a set…

Computation and Language · Computer Science 2023-07-10 Krishnapriya Vishnubhotla , Frank Rudzicz , Graeme Hirst , Adam Hammond