中文
相关论文

相关论文: V\=arta: A Large-Scale Headline-Generation Dataset…

200 篇论文

The research on text summarization for low-resource Indian languages has been limited due to the availability of relevant datasets. This paper presents a summary of various deep-learning approaches used for the ILSUM 2022 Indic language…

计算与语言 · 计算机科学 2022-12-13 Rahul Tangsali , Aabha Pingle , Aditya Vyawahare , Isha Joshi , Raviraj Joshi

We present the largest publicly available synthetic OCR benchmark dataset for Indic languages. The collection contains a total of 90k images and their ground truth for 23 Indic languages. OCR model validation in Indic languages require a…

计算机视觉与模式识别 · 计算机科学 2022-05-06 Naresh Saini , Promodh Pinto , Aravinth Bheemaraj , Deepak Kumar , Dhiraj Daga , Saurabh Yadav , Srihari Nagaraj

A diversity of tasks use language models trained on semantic similarity data. While there are a variety of datasets that capture semantic similarity, they are either constructed from modern web data or are relatively small datasets created…

计算与语言 · 计算机科学 2023-08-25 Emily Silcock , Melissa Dell

We present sentence aligned parallel corpora across 10 Indian Languages - Hindi, Telugu, Tamil, Malayalam, Gujarati, Urdu, Bengali, Oriya, Marathi, Punjabi, and English - many of which are categorized as low resource. The corpora are…

计算与语言 · 计算机科学 2020-07-16 Shashank Siripragada , Jerin Philip , Vinay P. Namboodiri , C V Jawahar

While there has been significant progress towards developing NLU resources for Indic languages, syntactic evaluation has been relatively less explored. Unlike English, Indic languages have rich morphosyntax, grammatical genders, free linear…

计算与语言 · 计算机科学 2021-10-05 Rajaswa Patil , Jasleen Dhillon , Siddhant Mahurkar , Saumitra Kulkarni , Manav Malhotra , Veeky Baths

Large Language Models (LLMs) have made significant progress in incorporating Indic languages within multilingual models. However, it is crucial to quantitatively assess whether these languages perform comparably to globally dominant ones,…

计算与语言 · 计算机科学 2024-10-31 Pritika Rohera , Chaitrali Ginimav , Akanksha Salunke , Gayatri Sawant , Raviraj Joshi

Being less resource languages, Indian-Indian and English-Indian language MT system developments faces the difficulty to translate various lexical phenomena. In this paper, we present our work on a comparative study of 440 phrase-based…

计算与语言 · 计算机科学 2017-10-09 Sreelekha S , Pushpak Bhattacharyya

Recent NLP advances focus primarily on standardized languages, leaving most low-resource dialects under-served especially in Indian scenarios. In India, the issue is particularly important: despite Hindi being the third most spoken language…

Language modeling has witnessed remarkable advancements in recent years, with Large Language Models (LLMs) like ChatGPT setting unparalleled benchmarks in human-like text generation. However, a prevailing limitation is the…

计算与语言 · 计算机科学 2023-11-13 Abhinand Balachandran

Babel Briefings is a novel dataset featuring 4.7 million news headlines from August 2020 to November 2021, across 30 languages and 54 locations worldwide with English translations of all articles included. Designed for natural language…

计算与语言 · 计算机科学 2024-03-29 Felix Leeb , Bernhard Schölkopf

Language documentation projects often involve the creation of annotated text in a format such as interlinear glossed text (IGT), which captures fine-grained morphosyntactic analyses in a morpheme-by-morpheme format. However, there are few…

计算与语言 · 计算机科学 2024-11-14 Michael Ginn , Lindia Tjuatja , Taiqi He , Enora Rice , Graham Neubig , Alexis Palmer , Lori Levin

We present VBART, the first Turkish sequence-to-sequence Large Language Models (LLMs) pre-trained on a large corpus from scratch. VBART are compact LLMs based on good ideas leveraged from BART and mBART models and come in two sizes, Large…

计算与语言 · 计算机科学 2024-03-15 Meliksah Turker , Mehmet Erdi Ari , Aydin Han

This paper introduces AfriHG -- a news headline generation dataset created by combining from XLSum and MasakhaNEWS datasets focusing on 16 languages widely spoken by Africa. We experimented with two seq2eq models (mT5-base and AfriTeVa V2),…

计算与语言 · 计算机科学 2024-12-31 Toyib Ogunremi , Serah Akojenu , Anthony Soronnadi , Olubayo Adekanmbi , David Ifeoluwa Adelani

The conversion of content from one language to another utilizing a computer system is known as Machine Translation (MT). Various techniques have come up to ensure effective translations that retain the contextual and lexical interpretation…

计算与语言 · 计算机科学 2024-01-15 Sudhansu Bala Das , Leo Raphael Rodrigues , Tapas Kumar Mishra , Bidyut Kr. Patra

The milestone improvements brought about by deep representation learning and pre-training techniques have led to large performance gains across downstream NLP, IR and Vision tasks. Multimodal modeling techniques aim to leverage large…

计算机视觉与模式识别 · 计算机科学 2023-02-21 Krishna Srinivasan , Karthik Raman , Jiecao Chen , Michael Bendersky , Marc Najork

Existing research on news summarization primarily focuses on single-language single-document (SLSD), single-language multi-document (SLMD) or cross-language single-document (CLSD). However, in real-world scenarios, news about a…

计算与语言 · 计算机科学 2024-10-15 Shengxiang Gao , Fang nan , Yongbing Zhang , Yuxin Huang , Kaiwen Tan , Zhengtao Yu

We present a collection of open, machine-readable document datasets covering parliamentary proceedings, legal judgments, government publications, news, and tourism statistics from Sri Lanka. The collection currently comprises of 269,194…

计算与语言 · 计算机科学 2026-05-18 Nuwan I. Senaratna

Generative AI models have shown impressive performance on many Natural Language Processing tasks such as language understanding, reasoning, and language generation. An important question being asked by the AI community today is about the…

Wordnets are rich lexico-semantic resources. Linked wordnets are extensions of wordnets, which link similar concepts in wordnets of different languages. Such resources are extremely useful in many Natural Language Processing (NLP)…

计算与语言 · 计算机科学 2022-01-11 Diptesh Kanojia , Kevin Patel , Pushpak Bhattacharyya