中文
相关论文

相关论文: Documenting Geographically and Contextually Divers…

200 篇论文

Multimodal large language models (MLLMs) enhance the capabilities of standard large language models by integrating and processing data from multiple modalities, including text, vision, audio, video, and 3D environments. Data plays a pivotal…

Recent regulatory initiatives like the European AI Act and relevant voices in the Machine Learning (ML) community stress the need to describe datasets along several key dimensions for trustworthy AI, such as the provenance processes and…

数字图书馆 · 计算机科学 2024-05-27 Joan Giner-Miguelez , Abel Gómez , Jordi Cabot

A recent study has shown that large-scale visual datasets are very biased: they can be easily classified by modern neural networks. However, the concrete forms of bias among these datasets remain unclear. In this study, we propose a…

计算机视觉与模式识别 · 计算机科学 2024-12-04 Boya Zeng , Yida Yin , Zhuang Liu

Information visualization and natural language are intricately linked. However, the majority of research and relevant work in information and data visualization (and human-computer interaction) involve English-speaking populations as both…

人机交互 · 计算机科学 2023-10-04 Noëlle Rakotondravony , Priya Dhawka , Melanie Bancilhon

Recent advances in large language models (LLMs) have led to new summarization strategies, offering an extensive toolkit for extracting important information. However, these approaches are frequently limited by their reliance on isolated…

人工智能 · 计算机科学 2024-06-21 Pranav Janjani , Mayank Palan , Sarvesh Shirude , Ninad Shegokar , Sunny Kumar , Faruk Kazi

Large language models are commonly trained on dominant languages like English, and their representation of low resource languages typically reflects cultural and linguistic biases present in the source language materials. Using the Serbian…

计算与语言 · 计算机科学 2025-12-12 Smiljana Antonijevic Ubois

Research data are often released upon journal publication to enable result verification and reproducibility. For that reason, research dissemination infrastructures typically support diverse datasets coming from numerous disciplines, from…

数字图书馆 · 计算机科学 2023-05-29 Ana Trisovic

In this position paper, we describe our perspective on how meaningful resources for lower-resourced languages should be developed in connection with the speakers of those languages. We first examine two massively multilingual resources in…

计算与语言 · 计算机科学 2022-02-25 Constantine Lignos , Nolan Holley , Chester Palen-Michel , Jonne Sälevä

Smart devices generate vast amounts of big data, mainly in the form of sensor data. While allowing for the prediction of many aspects of human behaviour (e.g., physical activities, transportation modes), this data has a major limitation in…

人机交互 · 计算机科学 2024-09-11 Fausto Giunchiglia , Xiaoyue Li

Arabic remains one of the most underrepresented languages in natural language processing research, particularly in medical applications, due to the limited availability of open-source data and benchmarks. The lack of resources hinders…

We present a dataset of 833k paragraphs extracted from CC-BY licensed scientific publications, classified into four categories: acknowledgments, data mentions, software/code mentions, and clinical trial mentions. The paragraphs are…

计算与语言 · 计算机科学 2025-10-28 Eric Jeangirard

Current AI models often fail to account for local context and language, given the predominance of English and Western internet content in their training data. This hinders the global relevance, usefulness, and safety of these models as they…

Within the past few decades we have witnessed digital revolution, which moved scholarly communication to electronic media and also resulted in a substantial increase in its volume. Nowadays keeping track with the latest scientific…

数字图书馆 · 计算机科学 2017-10-30 Dominika Tkaczyk

Data lakes have emerged as an alternative to data warehouses for the storage, exploration and analysis of big data. In a data lake, data are stored in a raw state and bear no explicit schema. Thence, an efficient metadata system is…

数据库 · 计算机科学 2019-05-13 Pegdwendé Sawadogo , Tokio Kibata , Jérôme Darmont

This study explores methods to increase data volume for low-resource languages using techniques such as crowdsourcing, pseudo-labeling, advanced data preprocessing and various permissive data sources such as audiobooks, Common Voice,…

In the rapidly evolving digital era, there is an increasing demand for concise information as individuals seek to distil key insights from various sources. Recent attention from researchers on Multi-document Summarisation (MDS) has resulted…

计算与语言 · 计算机科学 2024-09-19 Kushan Hewapathirana , Nisansa de Silva , C. D. Athuraliya

Large Language Models (LLMs) have demonstrated remarkable capabilities in important tasks such as natural language understanding and language generation, and thus have the potential to make a substantial impact on our society. Such…

计算与语言 · 计算机科学 2024-05-24 Zhongwei Wan , Xin Wang , Che Liu , Samiul Alam , Yu Zheng , Jiachen Liu , Zhongnan Qu , Shen Yan , Yi Zhu , Quanlu Zhang , Mosharaf Chowdhury , Mi Zhang

We present HowSumm, a novel large-scale dataset for the task of query-focused multi-document summarization (qMDS), which targets the use-case of generating actionable instructions from a set of sources. This use-case is different from the…

计算与语言 · 计算机科学 2021-10-12 Odellia Boni , Guy Feigenblat , Guy Lev , Michal Shmueli-Scheuer , Benjamin Sznajder , David Konopnicki

Despite advancements in conversational AI, language models encounter challenges to handle diverse conversational tasks, and existing dialogue dataset collections often lack diversity and comprehensiveness. To tackle these issues, we…

计算与语言 · 计算机科学 2024-02-06 Jianguo Zhang , Kun Qian , Zhiwei Liu , Shelby Heinecke , Rui Meng , Ye Liu , Zhou Yu , Huan Wang , Silvio Savarese , Caiming Xiong

Practically all large language models have been pre-trained on data that is subject to global uncertainty related to copyright infringement and breach of contract. This creates potential risk for users and developers due to this uncertain…

计算与语言 · 计算机科学 2025-04-11 Michael J Bommarito , Jillian Bommarito , Daniel Martin Katz