English
Related papers

Related papers: The Open Language Archives Community and Asian Lan…

200 papers

Digital Archive to MPEG-21 DIDL (D2D) analyzes the contents of the digital archive and produces an MPEG-21 Digital Item Declaration Language (DIDL) encapsulating the analysis results. DIDL is an extensible XML-based language that aggregates…

Digital Libraries · Computer Science 2007-05-23 Suchitra Manepalli , Giridhar Manepalli , Michael L. Nelson

Investigative journalism in recent years is confronted with two major challenges: 1) vast amounts of unstructured data originating from large text collections such as leaks or answers to Freedom of Information requests, and 2) multi-lingual…

Computation and Language · Computer Science 2018-07-17 Gregor Wiedemann , Seid Muhie Yimam , Chris Biemann

This paper accompanies the software documentation data set for machine translation, a parallel evaluation data set of data originating from the SAP Help Portal, that we released to the machine translation community for research purposes. It…

Computation and Language · Computer Science 2020-11-13 Bianka Buschbeck , Miriam Exel

In the era of large language models (LLMs), high-quality, domain-rich, and continuously evolving datasets capturing expert-level knowledge, core human values, and reasoning are increasingly valuable. This position paper argues that…

Computers and Society · Computer Science 2025-05-29 Hao Sun , Yunyi Shen , Mihaela van der Schaar

Numerous digital humanities projects maintain their data collections in the form of text, images, and metadata. While data may be stored in many formats, from plain text to XML to relational databases, the use of the resource description…

Digital Libraries · Computer Science 2014-06-03 Jakob Huber , Timo Sztyler , Jan Noessner , Jaimie Murdock , Colin Allen , Mathias Niepert

Through discovery of meso-scale structures, community detection methods contribute to the understanding of complex networks. Many community finding methods, however, rely on disjoint clustering techniques, in which node membership is…

Social and Information Networks · Computer Science 2022-11-23 Akhil Jakatdar , Baqiao Liu , Tandy Warnow , George Chacko

Advances in speech and language technologies enable tools such as voice-search, text-to-speech, speech recognition and machine translation. These are however only available for high resource languages like English, French or Chinese.…

A Language Model is a term that encompasses various types of models designed to understand and generate human communication. Large Language Models (LLMs) have gained significant attention due to their ability to process text with human-like…

Computation and Language · Computer Science 2024-06-12 Sylvio Barbon Junior , Paolo Ceravolo , Sven Groppe , Mustafa Jarrar , Samira Maghool , Florence Sèdes , Soror Sahri , Maurice Van Keulen

Low-resource languages serve as invaluable repositories of human history, embodying cultural evolution and intellectual diversity. Despite their significance, these languages face critical challenges, including data scarcity and…

The IVOA works towards standardising interoperability and curation of data and service holdings of the global astrophysical community. Within the IVOA, the Data Access Layer (DAL) Working Group's goal is to provide technical standards for…

Instrumentation and Methods for Astrophysics · Physics 2020-12-03 Marco Molinaro , James Dempsey

Linked Open Data (LOD) is the publicly available RDF data in the Web. Each LOD entity is identfied by a URI and accessible via HTTP. LOD encodes globalscale knowledge potentially available to any human as well as artificial intelligence…

AARC (Authentication and Authorisation for Research Communities) is a two-year EC-funded project to develop and pilot an integrated cross-discipline authentication and authorisation framework, building on existing authentication and…

With the rise of artificial intelligence (AI) and the growing use of deep-learning architectures, the question of ethics, transparency and fairness of AI systems has become a central concern within the research community. We address…

Computation and Language · Computer Science 2020-03-19 Mahault Garnerin , Solange Rossato , Laurent Besacier

Multimodal Large Language Models (mLLMs) are trained on a large amount of text-image data. While most mLLMs are trained on caption-like data only, Alayrac et al. (2022) showed that additionally training them on interleaved sequences of text…

Computation and Language · Computer Science 2025-05-30 Matthieu Futeral , Armel Zebaze , Pedro Ortiz Suarez , Julien Abadji , Rémi Lacroix , Cordelia Schmid , Rachel Bawden , Benoît Sagot

Open coding, a key inductive step in qualitative research, discovers and constructs concepts from human datasets. However, capturing extensive and nuanced aspects or "coding moments" can be challenging, especially with large discourse…

Computation and Language · Computer Science 2025-04-07 John Chen , Alexandros Lotsos , Grace Wang , Lexie Zhao , Bruce Sherin , Uri Wilensky , Michael Horn

We consider the problem of evaluating, and comparing computational policies in the Open Digital Rights Language (ODRL), which has become the de facto standard for governing the access and usage of digital resources. Although preliminary…

Artificial Intelligence · Computer Science 2025-09-09 Jaime Osvaldo Salas , Paolo Pareti , Semih Yumuşak , Soulmaz Gheisari , Luis-Daniel Ibáñez , George Konstantinidis

The goal of the present chapter is to explore the possibility of providing the research (but also the industrial) community that commonly uses spoken corpora with a stable portfolio of well-documented standardised formats that allow a high…

Computation and Language · Computer Science 2012-03-06 Laurent Romary , Andreas Witt

This paper presents LOLA, a massively multilingual large language model trained on more than 160 languages using a sparse Mixture-of-Experts Transformer architecture. Our architectural and implementation choices address the challenge of…

Darija Open Dataset (DODa) represents an open-source project aimed at enhancing Natural Language Processing capabilities for the Moroccan dialect, Darija. With approximately 100,000 entries, DODa stands as the largest collaborative project…

Computation and Language · Computer Science 2024-05-24 Aissam Outchakoucht , Hamza Es-Samaali

Conversational information seeking (CIS) has been recognized as a major emerging research area in information retrieval. Such research will require data and tools, to allow the implementation and study of conversational systems. This paper…

Information Retrieval · Computer Science 2019-12-20 Hamed Zamani , Nick Craswell