English
Related papers

Related papers: L3Cube-IndicNews: News-based Short Text and Long D…

200 papers

Despite progress in comment-aware multimodal and multilingual summarization for English and Chinese, research in Indian languages remains limited. This study addresses this gap by introducing COSMMIC, a pioneering comment-sensitive…

Computation and Language · Computer Science 2025-06-19 Raghvendra Kumar , S. A. Mohammed Salman , Aryan Sahu , Tridib Nandi , Pragathi Y. P. , Sriparna Saha , Jose G. Moreno

Existing research on news summarization primarily focuses on single-language single-document (SLSD), single-language multi-document (SLMD) or cross-language single-document (CLSD). However, in real-world scenarios, news about a…

Computation and Language · Computer Science 2024-10-15 Shengxiang Gao , Fang nan , Yongbing Zhang , Yuxin Huang , Kaiwen Tan , Zhengtao Yu

We present the IndicNLP corpus, a large-scale, general-domain corpus containing 2.7 billion words for 10 Indian languages from two language families. We share pre-trained word embeddings trained on these corpora. We create news article…

Computation and Language · Computer Science 2020-05-04 Anoop Kunchukuttan , Divyanshu Kakwani , Satish Golla , Gokul N. C. , Avik Bhattacharyya , Mitesh M. Khapra , Pratyush Kumar

Large-scale news corpora support a wide range of research in Computational Social Science and NLP, yet access remains constrained: commercial archives impose prohibitive costs and licensing restrictions, while open alternatives like Common…

Computation and Language · Computer Science 2026-05-19 Ruggero Marino Lazzaroni , Jana Lasser , Kirill Solovev

The research on text summarization for low-resource Indian languages has been limited due to the availability of relevant datasets. This paper presents a summary of various deep-learning approaches used for the ILSUM 2022 Indic language…

Computation and Language · Computer Science 2022-12-13 Rahul Tangsali , Aabha Pingle , Aditya Vyawahare , Isha Joshi , Raviraj Joshi

Observing the damages that can be done by the rapid propagation of fake news in various sectors like politics and finance, automatic identification of fake news using linguistic analysis has drawn the attention of the research community.…

Computation and Language · Computer Science 2020-04-21 Md Zobaer Hossain , Md Ashraful Rahman , Md Saiful Islam , Sudipta Kar

The research on code-mixed data is limited due to the unavailability of dedicated code-mixed datasets and pre-trained language models. In this work, we focus on the low-resource Indian language Marathi which lacks any prior work in…

Computation and Language · Computer Science 2023-07-21 Tanmay Chavan , Omkar Gokhale , Aditya Kane , Shantanu Patankar , Raviraj Joshi

Research in Natural Language Processing (NLP) has increasingly become important due to applications such as text classification, text mining, sentiment analysis, POS tagging, named entity recognition, textual entailment, and many others.…

Artificial Intelligence · Computer Science 2022-10-21 Istiak Ahmad , Fahad AlQurashi , Rashid Mehmood

While progress has been made in the domain of video-language understanding, current state-of-the-art algorithms are still limited in their ability to understand videos at high levels of abstraction, such as news-oriented videos.…

Computer Vision and Pattern Recognition · Computer Science 2024-01-24 Shih-Han Chou , Matthew Kowal , Yasmin Niknam , Diana Moyano , Shayaan Mehdi , Richard Pito , Cheng Zhang , Ian Knopke , Sedef Akinli Kocak , Leonid Sigal , Yalda Mohsenzadeh

Sentence representation from vanilla BERT models does not work well on sentence similarity tasks. Sentence-BERT models specifically trained on STS or NLI datasets are shown to provide state-of-the-art performance. However, building these…

Computation and Language · Computer Science 2022-11-23 Ananya Joshi , Aditi Kajale , Janhavi Gadre , Samruddhi Deode , Raviraj Joshi

News headline generation is a crucial task in increasing productivity for both the readers and producers of news. This task can easily be aided by automated News headline-generation models. However, the presence of irrelevant headlines in…

Computation and Language · Computer Science 2024-04-18 Gopichand Kanumolu , Lokesh Madasu , Nirmal Surange , Manish Shrivastava

Recent progress in text classification has been focused on high-resource languages such as English and Chinese. For low-resource languages, amongst them most African languages, the lack of well-annotated data and effective preprocessing, is…

Computation and Language · Computer Science 2020-10-26 Rubungo Andre Niyongabo , Hong Qu , Julia Kreutzer , Li Huang

The rapid progress in question-answering (QA) systems has predominantly benefited high-resource languages, leaving Indic languages largely underrepresented despite their vast native speaker base. In this paper, we present IndicSQuAD, a…

Computation and Language · Computer Science 2025-05-14 Sharvi Endait , Ruturaj Ghatage , Aditya Kulkarni , Rajlaxmi Patil , Raviraj Joshi

Local/Native South African languages are classified as low-resource languages. As such, it is essential to build the resources for these languages so that they can benefit from advances in the field of natural language processing. In this…

Computation and Language · Computer Science 2023-06-14 Andani Madodonga , Vukosi Marivate , Matthew Adendorff

Natural Language Generation (NLG) for non-English languages is hampered by the scarcity of datasets in these languages. In this paper, we present the IndicNLG Benchmark, a collection of datasets for benchmarking NLG for 11 Indic languages.…

Computation and Language · Computer Science 2022-10-28 Aman Kumar , Himani Shrotriya , Prachi Sahu , Raj Dabre , Ratish Puduppully , Anoop Kunchukuttan , Amogh Mishra , Mitesh M. Khapra , Pratyush Kumar

We present MahaSTS, a human-annotated Sentence Textual Similarity (STS) dataset for Marathi, along with MahaSBERT-STS-v2, a fine-tuned Sentence-BERT model optimized for regression-based similarity scoring. The MahaSTS dataset consists of…

Computation and Language · Computer Science 2025-09-01 Aishwarya Mirashi , Ananya Joshi , Raviraj Joshi

Understanding the writing frame of news articles is vital for addressing social issues, and thus has attracted notable attention in the fields of communication studies. Yet, assessing such news article frames remains a challenge due to the…

Computation and Language · Computer Science 2024-05-24 Xi Chen , Mattia Samory , Scott Hale , David Jurgens , Przemyslaw A. Grabowicz

Known by more than 1.5 billion people in the Indian subcontinent, Indic languages present unique challenges and opportunities for natural language processing (NLP) research due to their rich cultural heritage, linguistic diversity, and…

Computation and Language · Computer Science 2025-01-29 Sankalp KJ , Ashutosh Kumar , Laxmaan Balaji , Nikunj Kotecha , Vinija Jain , Aman Chadha , Sreyoshi Bhaduri

The multi-sentential long sequence textual data unfolds several interesting research directions pertaining to natural language processing and generation. Though we observe several high-quality long-sequence datasets for English and other…

Computation and Language · Computer Science 2023-02-24 Rahul Gupta , Vivek Srivastava , Mayank Singh

We present a novel approach to data preparation for developing multilingual Indic large language model. Our meticulous data acquisition spans open-source and proprietary sources, including Common Crawl, Indic books, news articles, and…