中文
相关论文

相关论文: Inter-Rater Agreement Study on Readability Assessm…

200 篇论文

Bengali is the seventh most spoken language on earth, yet considered a low-resource language in the field of natural language processing (NLP). Question answering over unstructured text is a challenging NLP task as it requires understanding…

Context: Software programs can be written in different but functionally equivalent ways. Even though previous research has compared specific formatting elements to find out which alternatives affect code legibility, seeing the bigger…

软件工程 · 计算机科学 2023-06-02 Delano Oliveira , Reydne Santos , Fernanda Madeiral , Hidehiko Masuhara , Fernando Castor

Despite the existence of numerous Optical Character Recognition (OCR) tools, the lack of comprehensive open-source systems hampers the progress of document digitization in various low-resource languages, including Bengali. Low-resource…

Based on the sense definition of words available in the Bengali WordNet, an attempt is made to classify the Bengali sentences automatically into different groups in accordance with their underlying senses. The input sentences are collected…

计算与语言 · 计算机科学 2015-08-07 Alok Ranjan Pal , Diganta Saha , Niladri Sekhar Dash

Despite its widespread use, Bengali lacks a robust automated International Phonetic Alphabet (IPA) transcription system that effectively supports both standard language and regional dialectal texts. Existing approaches struggle to handle…

计算与语言 · 计算机科学 2026-02-05 Jakir Hasan , Shrestha Datta , Md Saiful Islam , Shubhashis Roy Dipta , Ameya Debnath

Despite being the seventh most widely spoken language in the world, Bengali has received much less attention in machine translation literature due to being low in resources. Most publicly available parallel corpora for Bengali are not large…

计算与语言 · 计算机科学 2020-10-08 Tahmid Hasan , Abhik Bhattacharjee , Kazi Samin , Masum Hasan , Madhusudan Basak , M. Sohel Rahman , Rifat Shahriyar

Some users of social media are spreading racist, sexist, and otherwise hateful content. For the purpose of training a hate speech detection system, the reliability of the annotations is crucial, but there is no universally agreed-upon…

计算与语言 · 计算机科学 2017-01-30 Björn Ross , Michael Rist , Guillermo Carbonell , Benjamin Cabrera , Nils Kurowsky , Michael Wojatzki

Text classification has been one of the earliest problems in NLP. Over time the scope of application areas has broadened and the difficulty of dealing with new areas (e.g., noisy social media content) has increased. The problem-solving…

计算与语言 · 计算机科学 2020-11-10 Tanvirul Alam , Akib Khan , Firoj Alam

This research investigates the effectiveness of ChatGPT, an AI language model by OpenAI, in translating English into Hindi, Telugu, and Kannada languages, aimed at assisting tourists in India's linguistically diverse environment. To measure…

计算与语言 · 计算机科学 2023-07-31 Sanjana Kolar , Rohit Kumar

Several multilingual benchmark datasets have been developed in a semi-automatic manner in the recent past to measure progress and understand the state-of-the-art in the multilingual capabilities of Large Language Models. However, there is…

计算与语言 · 计算机科学 2025-11-14 Chalamalasetti Kranti , Gabriel Bernier-Colborne , Yvan Gauthier , Sowmya Vajjala

Handwritten Text Recognition (HTR) is a well-established research area. In contrast, Handwritten Text Generation (HTG) is an emerging field with significant potential. This task is challenging due to the variation in individual handwriting…

计算机视觉与模式识别 · 计算机科学 2025-12-29 Md. Rakibul Islam , Md. Kamrozzaman Bhuiyan , Safwan Muntasir , Arifur Rahman Jawad , Most. Sharmin Sultana Samu

This article presents the creation of an Estonian-language dataset for document-level subjectivity, analyzes the resulting annotations, and reports an initial experiment of automatic subjectivity analysis using a large language model (LLM).…

计算与语言 · 计算机科学 2025-12-11 Karl Gustav Gailit , Kadri Muischnek , Kairit Sirts

The sentiment analysis task in Tamil-English code-mixed texts has been explored using advanced transformer-based models. Challenges from grammatical inconsistencies, orthographic variations, and phonetic ambiguities have been addressed. The…

Researchers have traditionally recruited native speakers to provide annotations for widely used benchmark datasets. However, there are languages for which recruiting native speakers can be difficult, and it would help to find learners of…

计算与语言 · 计算机科学 2023-05-30 Haneul Yoo , Rifki Afina Putri , Changyoon Lee , Youngin Lee , So-Yeon Ahn , Dongyeop Kang , Alice Oh

Bengali is one of the most spoken languages in the world with over 300 million speakers globally. Despite its popularity, research into the development of Bengali speech recognition systems is hindered due to the lack of diverse open-source…

Observing the damages that can be done by the rapid propagation of fake news in various sectors like politics and finance, automatic identification of fake news using linguistic analysis has drawn the attention of the research community.…

计算与语言 · 计算机科学 2020-04-21 Md Zobaer Hossain , Md Ashraful Rahman , Md Saiful Islam , Sudipta Kar

A recitation is a way of combining the words together so that they have a sense of rhythm and thus an emotional content is imbibed within. In this study we envisaged to answer these questions in a scientific manner taking into consideration…

音频与语音处理 · 电气工程与系统科学 2020-08-06 Chirayata Bhattacharyya , Sourya Sengupta , Sayan Nag , Shankha Sanyal , Archi Banerjee , Ranjan Sengupta , Dipak Ghosh

One out of four children in India are leaving grade eight without basic reading skills. Measuring the reading levels in a vast country like India poses significant hurdles. Recent advances in machine learning opens up the possibility of…

音频与语音处理 · 电气工程与系统科学 2020-02-14 Dolly Agarwal , Jayant Gupchup , Nishant Baghel

This research presents a comprehensive investigation into Bangla authorship attribution, introducing a new balanced benchmark corpus BARD10 (Bangla Authorship Recognition Dataset of 10 authors) and systematically analyzing the impact of…

计算与语言 · 计算机科学 2025-11-12 Abdullah Muhammad Moosa , Nusrat Sultana , Mahdi Muhammad Moosa , Md. Miraiz Hossain

Human annotation remains the foundation of reliable and interpretable data in Natural Language Processing (NLP). As annotation and evaluation tasks continue to expand, from categorical labelling to segmentation, subjective judgment, and…

计算与语言 · 计算机科学 2026-04-02 Joseph James