English
Related papers

Related papers: The UN Security Council debates 1992-2023

200 papers

The majority of current Text-to-Speech (TTS) datasets, which are collections of individual utterances, contain few conversational aspects. In this paper, we introduce DailyTalk, a high-quality conversational speech dataset designed for…

Audio and Speech Processing · Electrical Eng. & Systems 2023-03-14 Keon Lee , Kyumin Park , Daeyoung Kim

Central banks around the world play a crucial role in maintaining economic stability. Deciphering policy implications in their communications is essential, especially as misinterpretations can disproportionately impact vulnerable…

Public knowledge of what is said in parliament is a tenet of democracy, and a critical resource for political science research. In Australia, following the British tradition, the written record of what is said in parliament is known as…

Digital Libraries · Computer Science 2023-09-25 Lindsay Katz , Rohan Alexander

This paper introduces the Ubuntu Dialogue Corpus, a dataset containing almost 1 million multi-turn dialogues, with a total of over 7 million utterances and 100 million words. This provides a unique resource for research into building…

Computation and Language · Computer Science 2016-07-26 Ryan Lowe , Nissan Pow , Iulian Serban , Joelle Pineau

Workplace meetings are vital to organizational collaboration, yet a large percentage of meetings are rated as ineffective. To help improve meeting effectiveness by understanding if the conversation is on topic, we create a comprehensive…

Computation and Language · Computer Science 2024-11-05 Yaran Fan , Jamie Pool , Senja Filipi , Ross Cutler

Code-switching (CS), the alternation between two or more languages within a single conversation, presents significant challenges for automatic speech recognition (ASR) systems. Existing Mandarin-English code-switching datasets often suffer…

Computation and Language · Computer Science 2025-03-13 Jiaming Zhou , Yujie Guo , Shiwan Zhao , Haoqin Sun , Hui Wang , Jiabei He , Aobo Kong , Shiyao Wang , Xi Yang , Yequan Wang , Yonghua Lin , Yong Qin

With its lenient moderation policies and long-standing associations with potentially unlawful activities, Telegram has become an incubator for problematic content, frequently featuring conspiratorial, hyper-partisan, and fringe narratives.…

Social and Information Networks · Computer Science 2024-11-01 Leonardo Blas , Luca Luceri , Emilio Ferrara

Compared to news and chat summarization, the development of meeting summarization is hugely decelerated by the limited data. To this end, we introduce a versatile Chinese meeting summarization dataset, dubbed VCSum, consisting of 239…

Computation and Language · Computer Science 2023-05-16 Han Wu , Mingjie Zhan , Haochen Tan , Zhaohui Hou , Ding Liang , Linqi Song

This project explores the application of Natural Language Processing (NLP) techniques to analyse United Nations General Assembly (UNGA) speeches. Using NLP allows for the efficient processing and analysis of large volumes of textual data,…

Computation and Language · Computer Science 2024-06-21 Mateusz Grzyb , Mateusz Krzyziński , Bartłomiej Sobieski , Mikołaj Spytek , Bartosz Pieliński , Daniel Dan , Anna Wróblewska

This document provides a brief description of the National Institute of Standards and Technology (NIST) speaker recognition evaluation (SRE) conversational telephone speech (CTS) Superset. The CTS Superset has been created in an attempt to…

Sound · Computer Science 2021-08-17 Seyed Omid Sadjadi

In this paper we present the ClaimBuster dataset of 23,533 statements extracted from all U.S. general election presidential debates and annotated by human coders. The ClaimBuster dataset can be leveraged in building computational methods to…

Computation and Language · Computer Science 2020-05-01 Fatma Arslan , Naeemul Hassan , Chengkai Li , Mark Tremayne

This paper presents the TikTok 2024 U.S. Presidential Election Dataset, a large-scale, resource designed to advance research into political communication and social media dynamics. The dataset comprises 3.14 million videos published on…

Social and Information Networks · Computer Science 2024-12-23 Gabriela Pinto , Charles Bickham , Tanishq Salkar , Joyston Menezes , Luca Luceri , Emilio Ferrara

The Speech Wikimedia Dataset is a publicly available compilation of audio with transcriptions extracted from Wikimedia Commons. It includes 1780 hours (195 GB) of CC-BY-SA licensed transcribed speech from a diverse set of scenarios and…

Artificial Intelligence · Computer Science 2023-08-31 Rafael Mosquera Gómez , Julián Eusse , Juan Ciro , Daniel Galvez , Ryan Hileman , Kurt Bollacker , David Kanter

Discord has evolved from a gaming-focused communication tool into a versatile platform supporting diverse online communities. Despite its large user base and active public servers, academic research on Discord remains limited due to data…

By evaluating Large Language Models (LLMs) through uniform, text-only interfaces, current academic benchmarks obscure how the unique designs and affordances of distinct commercial platforms shape real-world user behavior and system…

Computation and Language · Computer Science 2026-05-19 Yueru Yan , Tuc Nguyen , Bo Su , Melissa Lieffers , Thai Le

Automatic speech recognition (ASR) performs well for high-resource languages with abundant paired audio-transcript data, but its accuracy degrades sharply for most languages due to limited publicly available aligned data. To this end, we…

Computation and Language · Computer Science 2026-05-12 Antonis Asonitis , Luca A. Lanzendörfer , Frédéric Berdoz , Roger Wattenhofer

Hate speech represents a pervasive and detrimental form of online discourse, often manifested through an array of slurs, from hateful tweets to defamatory posts. As such speech proliferates, it connects people globally and poses significant…

Computation and Language · Computer Science 2025-05-06 Paloma Piot , Patricia Martín-Rodilla , Javier Parapar

Foreign policy analysis has been struggling to find ways to measure policy preferences and paradigm shifts in international political systems. This paper presents a novel, potential solution to this challenge, through the application of a…

Computation and Language · Computer Science 2017-07-13 Stefano Gurciullo , Slava Mikhaylov

Existing conversational datasets consist either of written proxies for dialog or small-scale transcriptions of natural speech. We introduce 'Interview': a large-scale (105K conversations) media dialog dataset collected from news interview…

Computation and Language · Computer Science 2020-04-08 Bodhisattwa Prasad Majumder , Shuyang Li , Jianmo Ni , Julian McAuley

Despite increasing awareness and research around fake news, there is still a significant need for datasets that specifically target racial slurs and biases within North American political speeches. This is particulary important in the…

Computation and Language · Computer Science 2024-01-09 Shaina Raza , Mizanur Rahman , Shardul Ghuge