English
Related papers

Related papers: The Open Language Archives Community and Asian Lan…

200 papers

The recent emergence and adoption of Machine Learning technology, and specifically of Large Language Models, has drawn attention to the need for systematic and transparent management of language data. This work proposes an approach to…

Datasets are foundational to many breakthroughs in modern artificial intelligence. Many recent achievements in the space of natural language processing (NLP) can be attributed to the finetuning of pre-trained models on a diverse set of…

In very recent years more attention has been placed on probing the role of pre-training data in Large Language Models (LLMs) downstream behaviour. Despite the importance, there is no public tool that supports such analysis of pre-training…

Computation and Language · Computer Science 2023-03-28 Thuy-Trang Vu , Xuanli He , Gholamreza Haffari , Ehsan Shareghi

Recent developments in Large Language Models (LLMs) have significantly expanded their applications across various domains. However, the effectiveness of LLMs is often constrained when operating individually in complex environments. This…

Artificial Intelligence · Computer Science 2024-05-08 Silvan Ferreira , Ivanovitch Silva , Allan Martins

With new emerging technologies, such as satellites and drones, archaeologists collect data over large areas. However, it becomes difficult to process such data in time. Archaeological data also have many different formats (images, texts,…

Databases · Computer Science 2021-07-26 Pengfei Liu , Sabine Loudcher , Jérôme Darmont , Camille Noûs

Traditional Online Public Access Catalogues (OPACs) are becoming less effective due to the rapid growth of scholarly literature. Conventional search methods, such as keyword indexing and Boolean queries, often fail to support efficient…

Digital Libraries · Computer Science 2026-04-03 M. S. Rajeevan , B. Mini Devi

The advancement of artificial intelligence (AI) hinges on the quality and accessibility of data, yet the current fragmentation and variability of data sources hinder efficient data utilization. The dispersion of data sources and diversity…

Digital Libraries · Computer Science 2024-07-22 Conghui He , Wei Li , Zhenjiang Jin , Chao Xu , Bin Wang , Dahua Lin

This project aims to optimize the linguistic indexing of the OpenAlex database by comparing the performance of various Python-based language identification procedures on different metadata corpora extracted from a manually-annotated article…

Computation and Language · Computer Science 2026-01-06 Maxime Holmberg Sainte-Marie , Diego Kozlowski , Lucía Céspedes , Vincent Larivière

Open Information Extraction (OIE) aims to extract objective structured knowledge from natural texts, which has attracted growing attention to build dedicated models with human experience. As the large language models (LLMs) have exhibited…

Computation and Language · Computer Science 2023-10-17 Ji Qi , Kaixuan Ji , Xiaozhi Wang , Jifan Yu , Kaisheng Zeng , Lei Hou , Juanzi Li , Bin Xu

We introduce the MultiLang Code Parser Dataset (MLCPD), a large-scale, language-agnostic dataset unifying syntactic and structural representations of code across ten major programming languages. MLCPD contains over seven million parsed…

Software Engineering · Computer Science 2025-10-21 Jugal Gajjar , Kamalasankari Subramaniakuppusamy

Most work in text classification and Natural Language Processing (NLP) focuses on English or a handful of other languages that have text corpora of hundreds of millions of words. This is creating a new version of the digital divide: the…

Computation and Language · Computer Science 2019-03-28 Meryem M'hamdi , Robert West , Andreea Hossmann , Michael Baeriswyl , Claudiu Musat

This paper analyses Conversational AI multi-agent interoperability frameworks and describes the novel architecture proposed by the Open Voice Interoperability initiative (Linux Foundation AI and DATA), also known briefly as OVON (Open Voice…

Artificial Intelligence · Computer Science 2024-07-30 Diego Gosmar , Deborah A. Dahl , Emmett Coin

OpenCitations is an infrastructure organization for open scholarship dedicated to the publication of open citation data as Linked Open Data using Semantic Web technologies, thereby providing a disruptive alternative to traditional…

Digital Libraries · Computer Science 2020-02-24 Silvio Peroni , David Shotton

Darija Open Dataset (DODa) is an open-source project for the Moroccan dialect. With more than 10,000 entries DODa is arguably the largest open-source collaborative project for Darija-English translation built for Natural Language Processing…

Computation and Language · Computer Science 2021-03-18 Aissam Outchakoucht , Hamza Es-Samaali

This study presents a novel framework for smart search in digital archival systems, leveraging the capabilities of Large Language Models (LLMs) to enhance information retrieval. By employing a Retrieval-Augmented Generation (RAG) approach,…

Artificial Intelligence · Computer Science 2025-01-14 Ha Dung Nguyen , Thi-Hoang Anh Nguyen , Thanh Binh Nguyen

In this position paper, we describe our perspective on how meaningful resources for lower-resourced languages should be developed in connection with the speakers of those languages. We first examine two massively multilingual resources in…

Computation and Language · Computer Science 2022-02-25 Constantine Lignos , Nolan Holley , Chester Palen-Michel , Jonne Sälevä

Sustainable Development Goals (SDGs) bring together the diverse development community and provide a clear set of development targets for 2030. Given a large number of actors and initiatives related to these goals, there is a need to have a…

Digital Libraries · Computer Science 2020-06-01 Lukas Pukelis , Nuria Bautista Puig , Mykola Skrynik , Vilius Stanciauskas

The global crisis of language endangerment meets a technological turning point as Generative AI (GenAI) and Large Language Models (LLMs) unlock new frontiers in automating corpus creation, transcription, translation, and tutoring. However,…

Computation and Language · Computer Science 2025-05-20 Vincent Koc

Recent advances in Natural Language Processing (NLP) have underscored the crucial role of high-quality datasets in building large language models (LLMs). However, while extensive resources and analyses exist for English, the landscape for…

Computation and Language · Computer Science 2025-10-16 Dasol Choi , Woomyoung Park , Youngsook Song

Human feedback on conversations with language language models (LLMs) is central to how these systems learn about the world, improve their capabilities, and are steered toward desirable and safe behaviors. However, this feedback is mostly…

‹ Prev 1 3 4 5 6 7 10 Next ›