English
Related papers

Related papers: Universal Dependency Treebank for Odia Language

200 papers

The Open Language Archives Community (OLAC) is an international partnership of institutions and individuals who are creating a worldwide virtual library of language resources. The Dublin Core (DC) Element Set and the OAI Protocol have…

Computation and Language · Computer Science 2007-05-23 Gary Simons , Steven Bird

This paper presents the process of building a neural machine translation system with support for English, Romanian, and Aromanian - an endangered Eastern Romance language. The primary contribution of this research is twofold: (1) the…

Computation and Language · Computer Science 2025-01-08 Alexandru-Iulius Jerpelea , Alina Rădoi , Sergiu Nisioi

We present the largest publicly available synthetic OCR benchmark dataset for Indic languages. The collection contains a total of 90k images and their ground truth for 23 Indic languages. OCR model validation in Indic languages require a…

Computer Vision and Pattern Recognition · Computer Science 2022-05-06 Naresh Saini , Promodh Pinto , Aravinth Bheemaraj , Deepak Kumar , Dhiraj Daga , Saurabh Yadav , Srihari Nagaraj

Opioid use disorder (OUD) is a chronic and relapsing condition that involves the continued and compulsive use of opioids despite harmful consequences. The development of medications with improved efficacy and safety profiles for OUD…

Biomolecules · Quantitative Biology 2023-03-02 Hongsong Feng , Jian Jiang , Guo-Wei Wei

Sentence simplification aims to make complex text more accessible by reducing linguistic complexity while preserving the original meaning. However, progress in this area remains limited for mid-resource and low-resource languages due to the…

This paper introduces UQA, a novel dataset for question answering and text comprehension in Urdu, a low-resource language with over 70 million native speakers. UQA is generated by translating the Stanford Question Answering Dataset…

Computation and Language · Computer Science 2024-07-24 Samee Arif , Sualeha Farid , Awais Athar , Agha Ali Raza

India's linguistic landscape is one of the most diverse in the world, comprising over 120 major languages and approximately 1,600 additional languages, with 22 officially recognized as scheduled languages in the Indian Constitution. Despite…

We propose a new contextual-compositional neural network layer that handles out-of-vocabulary (OOV) words in natural language processing (NLP) tagging tasks. This layer consists of a model that attends to both the character sequence and the…

Machine Learning · Computer Science 2019-12-17 Nicolas Garneau , Jean-Samuel Leboeuf , Yuval Pinter , Luc Lamontagne

We introduce UniRST, the first unified RST-style discourse parser capable of handling 18 treebanks in 11 languages without modifying their relation inventories. To overcome inventory incompatibilities, we propose and evaluate two training…

Computation and Language · Computer Science 2025-10-09 Elena Chistova

Large Language Models (LLMs) are now capable of generating text that closely resembles human writing, making them powerful tools for content creation, but this growing ability has also made it harder to tell whether a piece of text was…

Computation and Language · Computer Science 2025-10-21 Muhammad Ammar , Hadiya Murad Hadi , Usman Majeed Butt

Previous work has shown that isolated non-canonical sentences with Object-before-Subject (OSV) order are initially harder to process than their canonical counterparts with Subject-before-Object (SOV) order. Although this difficulty…

Computation and Language · Computer Science 2024-05-14 Sidharth Ranjan , Marten van Schijndel

Data-driven approaches for dependency parsing have been of great interest in Natural Language Processing for the past couple of decades. However, Sanskrit still lacks a robust purely data-driven dependency parser, probably with an exception…

Computation and Language · Computer Science 2020-04-20 Amrith Krishna , Ashim Gupta , Deepak Garasangi , Jivnesh Sandhan , Pavankumar Satuluri , Pawan Goyal

Word segmentation is a low-level NLP task that is non-trivial for a considerable number of languages. In this paper, we present a sequence tagging framework and apply it to word segmentation for a wide range of languages with different…

Computation and Language · Computer Science 2018-07-10 Yan Shao , Christian Hardmeier , Joakim Nivre

In this paper we present a sample treebank for Old English based on the UD Cairo sentences, collected and annotated as part of a classroom curriculum in Historical Linguistics. To collect the data, a sample of 20 sentences illustrating a…

Computation and Language · Computer Science 2025-06-13 Lauren Levine , Junghyun Min , Amir Zeldes

Dependency treebank is an important resource in any language. In this paper, we present our work on building BKTreebank, a dependency treebank for Vietnamese. Important points on designing POS tagset, dependency relations, and annotation…

Computation and Language · Computer Science 2018-02-22 Kiem-Hieu Nguyen

Information extraction suffers from its varying targets, heterogeneous structures, and demand-specific schemas. In this paper, we propose a unified text-to-structure generation framework, namely UIE, which can universally model different IE…

Computation and Language · Computer Science 2022-03-24 Yaojie Lu , Qing Liu , Dai Dai , Xinyan Xiao , Hongyu Lin , Xianpei Han , Le Sun , Hua Wu

To present the biodiversity information, a semantic model is required that connects all kinds of data about living creatures and their habitats. The model must be able to encode human knowledge for machines to be understood. Ontology offers…

Artificial Intelligence · Computer Science 2022-10-31 Archana Patel , Sarika Jain , Narayan C. Debnath , Vishal Lama

Data augmentation methods for neural machine translation are particularly useful when limited amount of training data is available, which is often the case when dealing with low-resource languages. We introduce a novel augmentation method,…

Computation and Language · Computer Science 2023-11-07 Attila Nagy , Dorina Lakatos , Botond Barta , Judit Ács

We present the IndicNLP corpus, a large-scale, general-domain corpus containing 2.7 billion words for 10 Indian languages from two language families. We share pre-trained word embeddings trained on these corpora. We create news article…

Computation and Language · Computer Science 2020-05-04 Anoop Kunchukuttan , Divyanshu Kakwani , Satish Golla , Gokul N. C. , Avik Bhattacharyya , Mitesh M. Khapra , Pratyush Kumar

In this paper, we present an approach to improve the accuracy of a strong transition-based dependency parser by exploiting dependency language models that are extracted from a large parsed corpus. We integrated a small number of features…

Computation and Language · Computer Science 2017-09-01 Juntao Yu , Bernd Bohnet