English
Related papers

Related papers: SiDiaC: Sinhala Diachronic Corpus

200 papers

Data-driven approaches for dependency parsing have been of great interest in Natural Language Processing for the past couple of decades. However, Sanskrit still lacks a robust purely data-driven dependency parser, probably with an exception…

Computation and Language · Computer Science 2020-04-20 Amrith Krishna , Ashim Gupta , Deepak Garasangi , Jivnesh Sandhan , Pavankumar Satuluri , Pawan Goyal

We present KS-PRET-5M, the largest publicly available pretraining dataset for the Kashmiri language, comprising 5,090,244 (5.09M) words, 27,692,959 (27.6M) characters, and a vocabulary of 295,433 (295.4K) unique word types. We assembled the…

Computation and Language · Computer Science 2026-04-14 Haq Nawaz Malik , Nahfid Nissar

This paper presents the "Speak & Improve Challenge 2025: Spoken Language Assessment and Feedback" -- a challenge associated with the ISCA SLaTE 2025 Workshop. The goal of the challenge is to advance research on spoken language assessment…

Computation and Language · Computer Science 2024-12-18 Mengjie Qian , Kate Knill , Stefano Banno , Siyuan Tang , Penny Karanasou , Mark J. F. Gales , Diane Nicholls

In NLP, text classification is one of the primary problems we try to solve and its uses in language analyses are indisputable. The lack of labeled training data made it harder to do these tasks in low resource languages like Amharic. The…

Computation and Language · Computer Science 2021-03-11 Israel Abebe Azime , Nebil Mohammed

Mining parallel document pairs for document-level machine translation (MT) remains challenging due to the limitations of existing Cross-Lingual Document Alignment (CLDA) techniques. Existing methods often rely on metadata such as URLs,…

Computation and Language · Computer Science 2025-11-11 Sanjay Suryanarayanan , Haiyue Song , Mohammed Safi Ur Rahman Khan , Anoop Kunchukuttan , Raj Dabre

Word meaning changes over time, depending on linguistic and extra-linguistic factors. Associating a word's correct meaning in its historical context is a central challenge in diachronic research, and is relevant to a range of NLP tasks,…

Computation and Language · Computer Science 2020-07-23 Valerio Perrone , Marco Palma , Simon Hengchen , Alessandro Vatri , Jim Q. Smith , Barbara McGillivray

Modern text-to-speech (TTS) systems use deep learning to synthesize speech increasingly approaching human quality, but they require a database of high quality audio-text sentence pairs for training. Malayalam, the official language of the…

Sound · Computer Science 2022-11-24 Deepa P Gopinath , Thennal D K , Vrinda V Nair , Swaraj K S , Sachin G

We present GS-BrainText, a curated dataset of 8,511 brain radiology reports from the Generation Scotland cohort, of which 2,431 are annotated for 24 brain disease phenotypes. This multi-site dataset spans five Scottish NHS health boards and…

Computation and Language · Computer Science 2026-03-30 Beatrice Alex , Claire Grover , Arlene Casey , Richard Tobin , Heather Whalley , William Whiteley

Learner corpus collects language data produced by L2 learners, that is second or foreign-language learners. This resource is of great relevance for second language acquisition research, foreign-language teaching, and automatic grammatical…

Computation and Language · Computer Science 2022-01-03 Yingying Wang , Cunliang Kong , Liner Yang , Yijun Wang , Xiaorong Lu , Renfen Hu , Shan He , Zhenghao Liu , Yun Chen , Erhong Yang , Maosong Sun

Modern Translation Systems heavily rely on high-quality, large parallel datasets for state-of-the-art performance. However, such resources are largely unavailable for most of the South Asian languages. Among them, Nepali and Tamang fall…

Recent years showed a strong increase in biomedical sciences and an inherent increase in publication volume. Extraction of specific information from these sources requires highly sophisticated text mining and information extraction tools.…

Computation and Language · Computer Science 2020-04-09 Johannes Kirschnick , Philippe Thomas , Roland Roller , Leonhard Hennig

Parallel text is required for building high-quality machine translation (MT) systems, as well as for other multilingual NLP applications. For many South Asian languages, such data is in short supply. In this paper, we described a new…

Computation and Language · Computer Science 2020-01-28 Barry Haddow , Faheem Kirefu

Language identification is used as the first step in many data collection and crawling efforts because it allows us to sort online text into language-specific buckets. However, many modern languages, such as Konkani, Kashmiri, Punjabi etc.,…

Computation and Language · Computer Science 2024-06-27 Milind Agarwal , Joshua Otten , Antonios Anastasopoulos

Character recognition techniques for printed documents are widely used for English language. However, the systems that are implemented to recognize Asian languages struggle to increase the accuracy of recognition. Among other Asian…

Computer Vision and Pattern Recognition · Computer Science 2014-12-25 G. I. Gunarathna , M. A. P. Chamikara , R. G. Ragel

Many Natural Language Processing and Computational Linguistics applications involves the generation of new texts based on some existing texts, such as summarization, text simplification and machine translation. However, there has been a…

Computation and Language · Computer Science 2018-04-12 Ping Chen , Fei Wu , Tong Wang , Wei Ding

Representing words and phrases into dense vectors of real numbers which encode semantic and syntactic properties is a vital constituent in natural language processing (NLP). The success of neural network (NN) models in NLP largely rely on…

Computation and Language · Computer Science 2021-01-01 Wazir Ali , Jay Kumar , Junyu Lu , Zenglin Xu

Despite having hundreds of millions of speakers, handwritten Devanagari text remains severely underrepresented in publicly available benchmark datasets. Existing resources are limited in scale, focus primarily on isolated characters or…

Computer Vision and Pattern Recognition · Computer Science 2026-02-23 Kunwar Arpit Singh , Ankush Prakash , Haroon R Lone

This paper describes additional aspects of a digital tool called the 'Textual History Tool'. We describe its various salient features with special reference to those of its features that may help the philologist digitize commentaries and…

Computation and Language · Computer Science 2022-01-06 Diptesh Kanojia , Malhar Kulkarni , Sayali Ghodekar , Eivind Kahrs , Pushpak Bhattacharyya

We propose a new database for information extraction from historical handwritten documents. The corpus includes 5,393 finding aids from six different series, dating from the 18th-20th centuries. Finding aids are handwritten documents that…

Computer Vision and Pattern Recognition · Computer Science 2023-04-27 Solène Tarride , Mélodie Boillet , Jean-François Moufflet , Christopher Kermorvant