English
Related papers

Related papers: The Open Language Archives Community and Asian Lan…

200 papers

A growing body of work shows that many problems in fairness, accountability, transparency, and ethics in machine learning systems are rooted in decisions surrounding the data collection and annotation process. In spite of its fundamental…

Machine Learning · Computer Science 2019-12-24 Eun Seo Jo , Timnit Gebru

This paper presents LAGOON -- an open source platform for understanding the complex ecosystems of Open Source Software (OSS) communities. The platform currently utilizes spatiotemporal graphs to store and investigate the artifacts produced…

Social and Information Networks · Computer Science 2022-01-28 Sourya Dey , Walt Woods

The Open Archive Initiative Protocol for Metadata Handling (OAI-PMHiii) is a standard that is seeing increased use as a means for exchanging structured metadata. OAI-PMH implementations must support Dublin Core as a metadata standard, with…

Information Retrieval · Computer Science 2011-01-04 Ranjeet Devarakonda , Giri Palanisamy , Bruce Wilson

Comprehensive Life Cycle Assessment (LCA) as a tool to account for the full range of environmental impacts of resource use in commodities or services is a first step in reducing these impacts. There is an increasing necessity to account for…

Physics and Society · Physics 2025-09-12 Hannah Wakeling , Kristin Lohwasser , Peter Millington

Southeast Asia (SEA) is a region rich in linguistic diversity and cultural variety, with over 1,300 indigenous languages and a population of 671 million people. However, prevailing AI models suffer from a significant lack of representation…

Large Language Models (LLMs) have emerged as powerful tools capable of understanding and generating human-like text, offering transformative potential across diverse domains. The Security Operations Center (SOC), responsible for…

Cryptography and Security · Computer Science 2025-09-23 Ali Habibzadeh , Farid Feyzi , Reza Ebrahimi Atani

The Data Web refers to the vast and rapidly increasing quantity of scientific, corporate, government and crowd-sourced data published in the form of Linked Open Data, which encourages the uniform representation of heterogeneous data items…

This paper discusses the requirements of current and emerging applications based on the Open Archives Initiative (OAI) and emphasizes the need for a common infrastructure to support them. Inspired by HTTP proxy, cache, gateway and web…

Digital Libraries · Computer Science 2007-05-23 Xiaoming Liu , Tim Brody , Stevan Harnad , Les Carr , Kurt Maly , Mohammad Zubair , Michael L. Nelson

This paper introduces a centralized, open-source dataset repository designed to advance NLP and NMT for Assamese, a low-resource language. The repository, available at GitHub, supports various tasks like sentiment analysis, named entity…

Computation and Language · Computer Science 2024-10-17 S. Tamang , D. J. Bora

While language technologies have advanced significantly, current approaches fail to address the complex sociocultural dimensions of linguistic preservation. AI Thinking proposes a meaning-centered framework that would transform…

Computation and Language · Computer Science 2025-02-24 Jose F Quesada

The increase in technological adoption worldwide comes with demands for novel tools to be used by the general population. Large Language Models (LLMs) provide a great opportunity in this respect, but their capabilities remain limited for…

Computation and Language · Computer Science 2025-10-13 Stefan Krsteski , Matea Tashkovska , Borjan Sazdov , Hristijan Gjoreski , Branislav Gerazov

This review paper provides a comprehensive overview of large language model (LLM) research directions within Indic languages. Indic languages are those spoken in the Indian subcontinent, including India, Pakistan, Bangladesh, Sri Lanka,…

Computation and Language · Computer Science 2024-06-17 Sankalp KJ , Vinija Jain , Sreyoshi Bhaduri , Tamoghna Roy , Aman Chadha

Consolidated access to current and reliable terms from different subject fields and languages is necessary for content creators and translators. Terminology is also needed in AI applications such as machine translation, speech recognition,…

Computation and Language · Computer Science 2022-07-15 Andis Lagzdiņš , Uldis Siliņš , Mārcis Pinnis , Toms Bergmanis , Artūrs Vasiļevskis , Andrejs Vasiļjevs

The paper reviews the hurdles while trying to implement the OLAC extension for Dravidian / Indian languages. The paper further explores the possibilities which could minimise or solve these problems. In this context, the Chinese system of…

Computation and Language · Computer Science 2009-09-08 B Prabhulla Chandran Pillai

Building LLMs for languages other than English is in great demand due to the unavailability and performance of multilingual LLMs, such as understanding the local context. The problem is critical for low-resource languages due to the need…

TalkBank is an online database that facilitates the sharing of linguistics research data. However, the existing TalkBank's API has limited data filtering and batch processing capabilities. To overcome these limitations, this paper…

Databases · Computer Science 2023-06-23 Man Ho Wong

The Open Annotation Core Data Model specifies an interoperable framework for creating associations between related resources, called annotations, using a methodology that conforms to the Architecture of the World Wide Web. Open Annotations…

Digital Libraries · Computer Science 2013-04-25 Robert Sanderson , Paolo Ciccarese , Herbert Van de Sompel

This paper presents an overview of a program designed to address the growing need for developing freely available speech resources for under-represented languages. At present we have released 38 datasets for building text-to-speech and…

The Archives Unleashed project aims to improve scholarly access to web archives through a multi-pronged strategy involving tool creation, process modeling, and community building - all proceeding concurrently in mutually-reinforcing…

Digital Libraries · Computer Science 2020-01-16 Nick Ruest , Jimmy Lin , Ian Milligan , Samantha Fritz

The Open Dataset of Audio Quality (ODAQ) was recently introduced to address the scarcity of openly available audio datasets with corresponding subjective quality scores. The dataset, released under permissive licenses, comprises audio…

Audio and Speech Processing · Electrical Eng. & Systems 2025-04-02 Sascha Dick , Christoph Thompson , Chih-Wei Wu , Matteo Torcoli , Pablo Delgado , Phillip A. Williams , Emanuel Habets