English
Related papers

Related papers: StoryDB: Broad Multi-language Narrative Dataset

200 papers

Collaborative stories, which are texts created through the collaborative efforts of multiple authors with different writing styles and intentions, pose unique challenges for NLP models. Understanding and generating such stories remains an…

Computation and Language · Computer Science 2023-05-16 Yulun Du , Lydia Chilton

Story video-text alignment, a core task in computational story understanding, aims to align video clips with corresponding sentences in their descriptions. However, progress on the task has been held back by the scarcity of manually…

Computation and Language · Computer Science 2024-10-04 Yidan Sun , Jianfei Yu , Boyang Li

The Speech Wikimedia Dataset is a publicly available compilation of audio with transcriptions extracted from Wikimedia Commons. It includes 1780 hours (195 GB) of CC-BY-SA licensed transcribed speech from a diverse set of scenarios and…

Artificial Intelligence · Computer Science 2023-08-31 Rafael Mosquera Gómez , Julián Eusse , Juan Ciro , Daniel Galvez , Ryan Hileman , Kurt Bollacker , David Kanter

We present the Multilingual Cloud Corpus, the first national-scale, parallel, multimodal linguistic dataset of Bangladesh's ethnic and indigenous languages. Despite being home to approximately 40 minority languages spanning four language…

Computation and Language · Computer Science 2026-03-09 Mohammad Mamun Or Rashid

While Large Language Models (LLMs) have significantly advanced Text-to-SQL performance, existing benchmarks predominantly focus on Western contexts and simplified schemas, leaving a gap in real-world, non-Western applications. We present…

Computation and Language · Computer Science 2026-04-16 Aviral Dawar , Roshan Karanth , Vikram Goyal , Dhruv Kumar

The emerging practice of data-driven storytelling is framing data using familiar narrative mechanisms such as slideshows, videos, and comics to make even highly complex phenomena understandable. However, current data stories are still not…

Human-Computer Interaction · Computer Science 2025-11-18 Zhenpeng Zhao , Niklas Elmqvist

The paper discusses the creation of a multimodal dataset of Russian-language scientific papers and testing of existing language models for the task of automatic text summarization. A feature of the dataset is its multimodal data, which…

Computation and Language · Computer Science 2024-05-14 Alena Tsanda , Elena Bruches

Spoken language translation has recently witnessed a resurgence in popularity, thanks to the development of end-to-end models and the creation of new corpora, such as Augmented LibriSpeech and MuST-C. Existing datasets involve language…

Computation and Language · Computer Science 2020-06-11 Changhan Wang , Juan Pino , Anne Wu , Jiatao Gu

We present a comprehensive evaluation of large language models for multilingual readability assessment. Existing evaluation resources lack domain and language diversity, limiting the ability for cross-domain and cross-lingual analyses. This…

Computation and Language · Computer Science 2024-10-17 Tarek Naous , Michael J. Ryan , Anton Lavrouk , Mohit Chandra , Wei Xu

We introduce ParaNames, a multilingual parallel name resource consisting of 118 million names spanning across 400 languages. Names are provided for 13.6 million entities which are mapped to standardized entity types (PER/LOC/ORG). Using…

Computation and Language · Computer Science 2022-07-13 Jonne Sälevä , Constantine Lignos

We present SpeakingFaces as a publicly-available large-scale multimodal dataset developed to support machine learning research in contexts that utilize a combination of thermal, visual, and audio data streams; examples include…

Human-Computer Interaction · Computer Science 2021-05-04 Madina Abdrakhmanova , Askat Kuzdeuov , Sheikh Jarju , Yerbolat Khassanov , Michael Lewis , Huseyin Atakan Varol

In this paper we share findings from our effort to build practical machine translation (MT) systems capable of translating across over one thousand languages. We describe results in three research domains: (i) Building clean, web-mined…

Analyzing literature involves tracking interactions between characters, locations, and themes. Visualization has the potential to facilitate the mapping and analysis of these complex relationships, but capturing structured information from…

Human-Computer Interaction · Computer Science 2025-08-12 Catherine Yeh , Tara Menon , Robin Singh Arya , Helen He , Moira Weigel , Fernanda Viégas , Martin Wattenberg

Scaling semantic parsing models for task-oriented dialog systems to new languages is often expensive and time-consuming due to the lack of available datasets. Available datasets suffer from several shortcomings: a) they contain few…

Computation and Language · Computer Science 2021-01-28 Haoran Li , Abhinav Arora , Shuohui Chen , Anchit Gupta , Sonal Gupta , Yashar Mehdad

The ability to transmit and receive complex information via language is unique to humans and is the basis of traditions, culture and versatile social interactions. Through the disruptive introduction of transformer based large language…

Computation and Language · Computer Science 2024-05-06 Patrick Krauss , Jannik Hösch , Claus Metzner , Andreas Maier , Peter Uhrig , Achim Schilling

Interpersonal language style shifting in dialogues is an interesting and almost instinctive ability of human. Understanding interpersonal relationship from language content is also a crucial step toward further understanding dialogues.…

Computation and Language · Computer Science 2020-12-07 Qi Jia , Hongru Huang , Kenny Q. Zhu

We present V\=arta, a large-scale multilingual dataset for headline generation in Indic languages. This dataset includes 41.8 million news articles in 14 different Indic languages (and English), which come from a variety of high-quality…

Computation and Language · Computer Science 2023-05-11 Rahul Aralikatte , Ziling Cheng , Sumanth Doddapaneni , Jackie Chi Kit Cheung

The advent of artificial intelligence (AI) has enabled a comprehensive exploration of materials for various applications. However, AI models often prioritize frequently encountered materials in the scientific literature, limiting the…

Materials Science · Physics 2023-08-29 Yang Jeong Park , Sung Eun Jerng , Jin-Sung Park , Choah Kwon , Chia-Wei Hsu , Zhichu Ren , Sungroh Yoon , Ju Li

In this paper we introduce the SchemaDB data-set; a collection of relational database schemata in both sql and graph formats. Databases are not commonly shared publicly for reasons of privacy and security, so schemata are not available for…

Databases · Computer Science 2025-05-30 Cody James Christopher , Kristen Moore , David Liebowitz

Clustering news across languages enables efficient media monitoring by aggregating articles from multilingual sources into coherent stories. Doing so in an online setting allows scalable processing of massive news streams. To this end, we…

Computation and Language · Computer Science 2018-09-05 Sebastião Miranda , Artūrs Znotiņš , Shay B. Cohen , Guntis Barzdins
‹ Prev 1 4 5 6 7 8 10 Next ›