English
Related papers

Related papers: WarCov -- Large multilabel and multimodal dataset …

200 papers

Measuring public attitudes toward wildlife provides crucial insights into our relationship with nature and helps monitor progress toward Global Biodiversity Framework targets. Yet, conducting such assessments at a global scale is…

Computation and Language · Computer Science 2024-05-06 Noah Giebink , Amrita Gupta , Diogo Verìssimo , Charlotte H. Chang , Tony Chang , Angela Brennan , Brett Dickson , Alex Bowmer , Jonathan Baillie

Large sense-annotated datasets are increasingly necessary for training deep supervised systems in Word Sense Disambiguation. However, gathering high-quality sense-annotated data for as many instances as possible is a laborious and expensive…

Computation and Language · Computer Science 2020-03-16 Tommaso Pasini , Jose Camacho-Collados

The COVID-19 epidemic is considered as the global health crisis of the whole society and the greatest challenge mankind faced since World War Two. Unfortunately, the fake news about COVID-19 is spreading as fast as the virus itself. The…

Social and Information Networks · Computer Science 2020-11-24 Yichuan Li , Bohan Jiang , Kai Shu , Huan Liu

In recent years, a series of Transformer-based models unlocked major improvements in general natural language understanding (NLU) tasks. Such a fast pace of research would not be possible without general NLU benchmarks, which allow for a…

Computation and Language · Computer Science 2020-05-05 Piotr Rybak , Robert Mroczkowski , Janusz Tracz , Ireneusz Gawlik

The COVID-19 infodemic, characterized by the rapid spread of misinformation and unverified claims related to the pandemic, presents a significant challenge. This paper presents a comparative analysis of the COVID-19 infodemic in the English…

Social and Information Networks · Computer Science 2023-11-15 Jia Luo , Daiyun Peng , Lei Shi , Didier El Baz , Xinran Liu

Following recent policy changes by X (Twitter) and other social media platforms, user interaction data has become increasingly difficult to access. These restrictions are impeding robust research pertaining to social and political phenomena…

Social and Information Networks · Computer Science 2024-07-01 Gabriele Di Bona , Emma Fraxanet , Björn Komander , Andrea Lo Sasso , Virginia Morini , Antoine Vendeville , Max Falkenberg , Alessandro Galeazzi

We introduce MULTI-EURLEX, a new multilingual dataset for topic classification of legal documents. The dataset comprises 65k European Union (EU) laws, officially translated in 23 languages, annotated with multiple labels from the EUROVOC…

Computation and Language · Computer Science 2021-09-08 Ilias Chalkidis , Manos Fergadiotis , Ion Androutsopoulos

Labeling training datasets has become a key barrier to building medical machine learning models. One strategy is to generate training labels programmatically, for example by applying natural language processing pipelines to text reports…

With the widespread application of artificial intelligence (AI), particularly deep learning (DL) and vision large language models (VLLMs), in skin disease diagnosis, the need for interpretability becomes crucial. However, existing…

Computer Vision and Pattern Recognition · Computer Science 2025-11-11 Yuhao Shen , Liyuan Sun , Yan Xu , Wenbin Liu , Shuping Zhang , Shawn Afvari , Zhongyi Han , Jiaoyan Song , Yongzhi Ji , Tao Lu , Xiaonan He , Xin Gao , Juexiao Zhou

Despite tremendous progress in computer vision, there has not been an attempt for machine learning on very large-scale medical image databases. We present an interleaved text/image deep learning system to extract and mine the semantic…

Computer Vision and Pattern Recognition · Computer Science 2015-05-05 Hoo-Chang Shin , Le Lu , Lauren Kim , Ari Seff , Jianhua Yao , Ronald M. Summers

Along with the COVID-19 pandemic, an "infodemic" of false and misleading information has emerged and has complicated the COVID-19 response efforts. Social networking sites such as Facebook and Twitter have contributed largely to the spread…

Computation and Language · Computer Science 2021-05-10 Mohamed Seghir Hadj Ameur , Hassina Aliane

Suicide is an important but often misunderstood problem, one that researchers are now seeking to better understand through social media. Due in large part to the fuzzy nature of what constitutes suicidal risks, most supervised approaches…

Machine Learning · Computer Science 2017-02-01 Tong Liu , Qijin Cheng , Christopher M. Homan , Vincent M. B. Silenzio

The paper describes the open Russian medical language understanding benchmark covering several task types (classification, question answering, natural language inference, named entity recognition) on a number of novel text sets. Given the…

Computation and Language · Computer Science 2022-07-14 Pavel Blinov , Arina Reshetnikova , Aleksandr Nesterov , Galina Zubkova , Vladimir Kokh

The paper discusses the creation of a multimodal dataset of Russian-language scientific papers and testing of existing language models for the task of automatic text summarization. A feature of the dataset is its multimodal data, which…

Computation and Language · Computer Science 2024-05-14 Alena Tsanda , Elena Bruches

The task of toxicity detection is still a relevant task, especially in the context of safe and fair LMs development. Nevertheless, labeled binary toxicity classification corpora are not available for all languages, which is understandable…

Computation and Language · Computer Science 2024-04-30 Daryna Dementieva , Valeriia Khylenko , Nikolay Babakov , Georg Groh

Contextual large language model embeddings are increasingly utilized for topic modeling and clustering. However, current methods often scale poorly, rely on opaque similarity metrics, and struggle in multilingual settings. In this work, we…

Computation and Language · Computer Science 2025-06-03 Hans W. A. Hanley , Zakir Durumeric

Data plays a vital role in machine learning studies. In the research of recommendation, both user behaviors and side information are helpful to model users. So, large-scale real scenario datasets with abundant user behaviors will contribute…

Information Retrieval · Computer Science 2021-06-14 Bin Hao , Min Zhang , Weizhi Ma , Shaoyun Shi , Xinxing Yu , Houzhi Shan , Yiqun Liu , Shaoping Ma

Music representation learning is central to music information retrieval and generation. While recent advances in multimodal learning have improved alignment between text and audio for tasks such as cross-modal music retrieval, text-to-music…

The healthcare environment is commonly referred to as "information-rich" but also "knowledge poor". Healthcare systems collect huge amounts of data from various sources: lab reports, medical letters, logs of medical tools or programs,…

Computation and Language · Computer Science 2024-01-22 Elena-Simona Apostol , Ciprian-Octavian Truică

Exabytes of data are generated daily by humans, leading to the growing need for new efforts in dealing with the grand challenges for multi-label learning brought by big data. For example, extreme multi-label classification is an active and…

Machine Learning · Computer Science 2021-11-18 Weiwei Liu , Haobo Wang , Xiaobo Shen , Ivor W. Tsang