English
Related papers

Related papers: The UN Security Council debates 1992-2023

200 papers

Reading comprehension by machine has been widely studied, but machine comprehension of spoken content is still a less investigated problem. In this paper, we release Open-Domain Spoken Question Answering Dataset (ODSQA) with more than three…

Computation and Language · Computer Science 2018-08-08 Chia-Hsuan Lee , Shang-Ming Wang , Huan-Cheng Chang , Hung-Yi Lee

The exchange of personal information in digital environments poses significant risks, including identity theft, privacy breaches, and data misuse. Addressing these challenges requires a deep understanding of user behavior and mental models…

Human-Computer Interaction · Computer Science 2025-09-22 Reza Shahriari , Eric D. Ragan

Group deliberation enables people to collaborate and solve problems, however, it is understudied due to a lack of resources. To this end, we introduce the first publicly available dataset containing collaborative conversations on solving a…

Computation and Language · Computer Science 2023-04-18 Georgi Karadzhov , Tom Stafford , Andreas Vlachos

When humans converse, what a speaker will say next significantly depends on what he sees. Unfortunately, existing dialogue models generate dialogue utterances only based on preceding textual contexts, and visual contexts are rarely…

Computation and Language · Computer Science 2021-06-01 Yuxian Meng , Shuhe Wang , Qinghong Han , Xiaofei Sun , Fei Wu , Rui Yan , Jiwei Li

Recent progress in speech processing has highlighted that high-quality performance across languages requires substantial training data for each individual language. While existing multilingual datasets cover many languages, they often…

Computation and Language · Computer Science 2025-10-28 Samuel Pfisterer , Florian Grötschla , Luca A. Lanzendörfer , Florian Yan , Roger Wattenhofer

We present a novel corpus of 445 human- and computer-generated documents, comprising about 27,000 clauses, annotated for semantic clause types and coherence relations that allow for nuanced comparison of artificial and natural discourse…

Enhancing user engagement through personalization in conversational agents has gained significance, especially with the advent of large language models that generate fluent responses. Personalized dialogue generation, however, is…

Computation and Language · Computer Science 2024-07-30 Yi-Pei Chen , Noriki Nishida , Hideki Nakayama , Yuji Matsumoto

We present a collection of open, machine-readable document datasets covering parliamentary proceedings, legal judgments, government publications, news, and tourism statistics from Sri Lanka. The collection currently comprises of 269,194…

Computation and Language · Computer Science 2026-05-18 Nuwan I. Senaratna

Local government meetings are the most common formal channel through which residents speak directly with elected officials, contest policies, and shape local agendas. However, data constraints typically limit the empirical study of these…

Econometrics · Economics 2026-04-24 Olivia Martin , Amar Venugopal

Large, openly licensed speech datasets are essential for building automatic speech recognition (ASR) systems, yet many widely spoken languages remain underrepresented in public resources. Pashto, spoken by more than 60 million people, has…

Computation and Language · Computer Science 2026-02-17 Jandad Jahani , Mursal Dawodi , Jawid Ahmad Baktash

We present spINAch, a large diachronic corpus of French speech from radio and television archives, balanced by speakers' gender, age (20-95 years old), and spanning 60 years from 1955 to 2015. The dataset includes over 320 hours of…

As the COVID-19 pandemic continues its march around the world, an unprecedented amount of open data is being generated for genetics and epidemiological research. The unparalleled rate at which many research groups around the world are…

Social and Information Networks · Computer Science 2021-08-10 Juan M. Banda , Ramya Tekumalla , Guanyu Wang , Jingyuan Yu , Tuo Liu , Yuning Ding , Katya Artemova , Elena Tutubalina , Gerardo Chowell

This paper was prepared for Open-Access-only publication as a guide reporting on Education (all aspects of space science and technology), Teaching (remote sensing and GIS, satellite meteorology and global climate, satellite communication,…

Physics Education · Physics 2025-07-24 Hans J. Haubold , Arak M. Mathai

Argumentation generation has attracted substantial research interest due to its central role in human reasoning and decision-making. However, most existing argumentative corpora focus on non-interactive, single-turn settings, either…

Computation and Language · Computer Science 2026-01-13 Yongkang Liu , Jiayang Yu , Mingyang Wang , Yiqun Zhang , Ercong Nie , Shi Feng , Daling Wang , Kaisong Song , Hinrich Schütze

In this paper, we present AISHELL-4, a sizable real-recorded Mandarin speech dataset collected by 8-channel circular microphone array for speech processing in conference scenario. The dataset consists of 211 recorded meeting sessions, each…

Sound · Computer Science 2021-08-11 Yihui Fu , Luyao Cheng , Shubo Lv , Yukai Jv , Yuxiang Kong , Zhuo Chen , Yanxin Hu , Lei Xie , Jian Wu , Hui Bu , Xin Xu , Jun Du , Jingdong Chen

With the increasing usage of the internet, more and more data is being digitized including parliamentary debates but they are in an unstructured format. There is a need to convert them into a structured format for linguistic analysis. Much…

Computation and Language · Computer Science 2018-08-22 Sakala Venkata Krishna Rohit , Navjyoti Singh

Building on current work on multilingual hate speech (e.g., Ousidhoum et al. (2019)) and hate speech reduction (e.g., Sap et al. (2020)), we present XTREMESPEECH, a new hate speech dataset containing 20,297 social media passages from…

Computation and Language · Computer Science 2022-03-23 Antonis Maronikolakis , Axel Wisiorek , Leah Nann , Haris Jabbar , Sahana Udupa , Hinrich Schuetze

Speech datasets are crucial for training Speech Language Technologies (SLT); however, the lack of diversity of the underlying training data can lead to serious limitations in building equitable and robust SLT products, especially along…

Speaker recognition is a widely used voice-based biometric technology with applications in various industries, including banking, education, recruitment, immigration, law enforcement, healthcare, and well-being. However, while dataset…

Computers and Society · Computer Science 2023-08-21 Casandra Rusti , Anna Leschanowsky , Carolyn Quinlan , Michaela Pnacek , Lauriane Gorce , Wiebke Hutiri

Existing argumentation datasets have succeeded in allowing researchers to develop computational methods for analyzing the content, structure and linguistic features of argumentative text. They have been much less successful in fostering…

Computation and Language · Computer Science 2019-09-26 Esin Durmus , Claire Cardie
‹ Prev 1 4 5 6 7 8 10 Next ›