中文
相关论文

相关论文: A Supervised Authorship Attribution Framework for …

200 篇论文

This research presents a comprehensive investigation into Bangla authorship attribution, introducing a new balanced benchmark corpus BARD10 (Bangla Authorship Recognition Dataset of 10 authors) and systematically analyzing the impact of…

计算与语言 · 计算机科学 2025-11-12 Abdullah Muhammad Moosa , Nusrat Sultana , Mahdi Muhammad Moosa , Md. Miraiz Hossain

Automatic Speech Recognition (ASR) for Bengali, the world's fifth most spoken language, remains a significant challenge, critically hindering technological accessibility for its over 270 million speakers. This challenge is compounded by two…

声音 · 计算机科学 2025-09-03 Swadhin Biswas , Imran , Tuhin Sheikh

In this paper we propose a learning paradigm for the problem of understanding spoken language. The basis of the work is in a formalization of the understanding problem as a communication problem. This results in the definition of a…

cmp-lg · 计算机科学 2008-02-03 Roberto Pieraccini , Esther Levin

Question-answering systems for Bengali have seen limited development, particularly in domain-specific applications. Leveraging advancements in natural language processing, this paper explores a fine-tuned BERT-Bangla model to address this…

计算与语言 · 计算机科学 2024-10-08 Subal Chandra Roy , Md Motaleb Hossen Manik

Lack of diverse perspectives causes neutrality bias in Wikipedia content leading to millions of worldwide readers getting exposed by potentially inaccurate information. Hence, neutrality bias detection and mitigation is a critical problem.…

计算与语言 · 计算机科学 2023-12-27 Ankita Maity , Anubhav Sharma , Rudra Dhar , Tushar Abhishek , Manish Gupta , Vasudeva Varma

The domain of Natural Language Processing (NLP) has experienced notable progress in the evolution of Bangla Question Answering (QA) systems. This paper presents a comprehensive review of seven research articles that contribute to the…

Maintaining anonymity in natural language communication remains a challenging task. Even when the number of candidate authors is large, standard authorship attribution techniques that analyze writing style predict the original author with…

计算与语言 · 计算机科学 2026-03-04 Haining Wang , Patrick Juola , Allen Riddell

This paper presents a novel approach to constructing an English-to-Telugu translation model by leveraging transfer learning techniques and addressing the challenges associated with low-resource languages. Utilizing the Bharat Parallel…

计算与语言 · 计算机科学 2025-04-09 Abhiram Reddy Yanampally

Author name disambiguation in bibliographic databases is the problem of grouping together scientific publications written by the same person, accounting for potential homonyms and/or synonyms. Among solutions to this problem, digital…

数字图书馆 · 计算机科学 2016-05-05 Gilles Louppe , Hussein Al-Natsheh , Mateusz Susik , Eamonn Maguire

Authorship attribution models fine-tuned with the same pretrained encoder, data, and loss can differ four-fold in performance depending only on their scoring mechanism. We use mechanistic interpretability tools to explain this gap.…

计算与语言 · 计算机科学 2026-05-27 Francis Kulumba , Guillaume Vimont , Laurent Romary , Florian Cafiero

Developing an automatic signature verification system is challenging and demands a large number of training samples. This is why synthetic handwriting generation is an emerging topic in document image analysis. Some handwriting synthesizers…

Document categorization is a technique where the category of a document is determined. In this paper three well-known supervised learning techniques which are Support Vector Machine(SVM), Na\"ive Bayes(NB) and Stochastic Gradient…

计算与语言 · 计算机科学 2017-01-31 Md. Saiful Islam , Fazla Elahi Md Jubayer , Syed Ikhtiar Ahmed

Due to the huge availability of documents in digital form, and the deception possibility raise bound to the essence of digital documents and the way they are spread, the authorship attribution problem has constantly increased its relevance.…

神经与进化计算 · 计算机科学 2014-10-01 Christian Napoli , Giuseppe Pappalardo , Emiliano Tramontana

Efforts on the research and development of OCR systems for Low-Resource Languages are relatively new. Low-resource languages have little training data available for training Machine Translation systems or other systems. Even though a vast…

计算与语言 · 计算机科学 2024-04-04 S M Rakib Hasan , Aakar Dhakal , Md Humaion Kabir Mehedi , Annajiat Alim Rasel

The Bangla language is the seventh most spoken language, with 265 million native and non-native speakers worldwide. However, English is the predominant language for online resources and technical knowledge, journals, and documentation.…

In this work, we have introduced Gaussian Smoothen Semantic Features (GSSF) for Better Semantic Selection for Indian regional language-based image captioning and introduced a procedure where we used the existing translation and English…

计算与语言 · 计算机科学 2020-02-18 Chiranjib Sur

Despite being the seventh most widely spoken language in the world, Bengali has received much less attention in machine translation literature due to being low in resources. Most publicly available parallel corpora for Bengali are not large…

计算与语言 · 计算机科学 2020-10-08 Tahmid Hasan , Abhik Bhattacharjee , Kazi Samin , Masum Hasan , Madhusudan Basak , M. Sohel Rahman , Rifat Shahriyar

Language documentation is inherently a time-intensive process; transcription, glossing, and corpus management consume a significant portion of documentary linguists' work. Advances in natural language processing can help to accelerate this…

计算与语言 · 计算机科学 2018-12-14 Graham Neubig , Patrick Littell , Chian-Yu Chen , Jean Lee , Zirui Li , Yu-Hsiang Lin , Yuyan Zhang

The wide accessibility of social media has provided linguistically under-represented communities with an extraordinary opportunity to create content in their native languages. This, however, comes with certain challenges in script…

计算与语言 · 计算机科学 2023-05-29 Sina Ahmadi , Antonios Anastasopoulos

For language documentation initiatives, transcription is an expensive resource: one minute of audio is estimated to take one hour and a half on average of a linguist's work (Austin and Sallabank, 2013). Recently, collecting aligned…

计算与语言 · 计算机科学 2019-10-14 Marcely Zanon Boito , Aline Villavicencio , Laurent Besacier