中文
相关论文

相关论文: StoryDB: Broad Multi-language Narrative Dataset

200 篇论文

With recent advancements in diffusion models, users can generate high-quality images by writing text prompts in natural language. However, generating images with desired details requires proper prompts, and it is often unclear how a model…

计算机视觉与模式识别 · 计算机科学 2023-07-07 Zijie J. Wang , Evan Montoya , David Munechika , Haoyang Yang , Benjamin Hoover , Duen Horng Chau

Work on shallow discourse parsing in English has focused on the Wall Street Journal corpus, the only large-scale dataset for the language in the PDTB framework. However, the data is not openly available, is restricted to the news domain,…

Timeline generation is of great significance for a comprehensive understanding of the development of events over time. Its goal is to organize news chronologically, which helps to identify patterns and trends that may be obscured when…

信息检索 · 计算机科学 2025-02-12 Xiaochen Liu , Yanan Zhang

Confidently making progress on multilingual modeling requires challenging, trustworthy evaluations. We present TyDi QA---a question answering dataset covering 11 typologically diverse languages with 204K question-answer pairs. The languages…

Data storytelling has seen rapid growth through a proliferation of examples, as well as theoretical and technical advancements contributed across multiple disciplines. In this paper, we present a comprehensive survey of data storytelling…

人机交互 · 计算机科学 2026-01-21 Leni Yang , Zezhong Wang , Xingyu Lan

Recent generative models have demonstrated impressive capabilities in generating realistic and visually pleasing images grounded on textual prompts. Nevertheless, a significant challenge remains in applying these models for the more…

计算机视觉与模式识别 · 计算机科学 2025-04-28 Xiaoqian Shen , Mohamed Elhoseiny

We describe an LSTM-based model which we call Byte-to-Span (BTS) that reads text as bytes and outputs span annotations of the form [start, length, label] where start positions, lengths, and labels are separate entries in our vocabulary.…

计算与语言 · 计算机科学 2016-04-05 Dan Gillick , Cliff Brunk , Oriol Vinyals , Amarnag Subramanya

Multilingual language models have significantly advanced due to rapid progress in natural language processing. Models like BLOOM 1.7B, trained on diverse multilingual datasets, aim to bridge linguistic gaps. However, their effectiveness in…

The increasing adoption of text-to-speech technologies has led to a growing demand for natural and emotive voices that adapt to a conversation's context and emotional tone. The Emotive Narrative Storytelling (EMNS) corpus is a unique speech…

计算与语言 · 计算机科学 2023-05-26 Kari Ali Noriy , Xiaosong Yang , Jian Jun Zhang

We present the Newspaper Bias Dataset (NewB), a text corpus of more than 200,000 sentences from eleven news sources regarding Donald Trump. While previous datasets have labeled sentences as either liberal or conservative, NewB covers the…

计算与语言 · 计算机科学 2023-09-12 Jerry Wei

While numerous architectures for long-range language models (LRLMs) have recently been proposed, a meaningful evaluation of their discourse-level language understanding capabilities has not yet followed. To this end, we introduce…

计算与语言 · 计算机科学 2022-04-26 Simeng Sun , Katherine Thai , Mohit Iyyer

We release 70,509 high-quality social networks extracted from multilingual fiction and nonfiction narratives. We additionally provide metadata for $\sim$30,000 of these texts (73\% nonfiction and 27\% fiction) written between 1800 and 1999…

计算与语言 · 计算机科学 2025-04-01 Sil Hamilton , Rebecca M. M. Hicke , David Mimno , Matthew Wilkens

Computational narrative understanding studies the identification, description, and interaction of the elements of a narrative: characters, attributes, events, and relations. Narrative research has given considerable attention to defining…

计算与语言 · 计算机科学 2025-04-22 Sabyasachee Baruah , Shrikanth Narayanan

Multilingual NLP often relies on dataset counts from centralized catalogues to characterize which languages are resource-rich or resource-poor. However, these catalogues record only one layer of dataset visibility: what has been registered…

计算与语言 · 计算机科学 2026-05-19 Zhiyin Tan , Changxu Duan

This paper introduces a large-scale multimodal and multilingual dataset that aims to facilitate research on grounding words to images in their contextual usage in language. The dataset consists of images selected to unambiguously illustrate…

计算与语言 · 计算机科学 2022-06-20 Josiah Wang , Pranava Madhyastha , Josiel Figueiredo , Chiraag Lala , Lucia Specia

Deepfakes represent a growing concern across domains such as disinformation, fraud, and non-consensual media. In particular, the rise of video conference and identity-driven attacks in high-stakes scenarios--such as impostor hiring--demands…

计算机视觉与模式识别 · 计算机科学 2026-04-08 Sarah Barrington , Maty Bohacek , Hany Farid

We train one multilingual model for dependency parsing and use it to parse sentences in several languages. The parsing model uses (i) multilingual word clusters and embeddings; (ii) token-level language information; and (iii)…

计算与语言 · 计算机科学 2016-07-27 Waleed Ammar , George Mulcaire , Miguel Ballesteros , Chris Dyer , Noah A. Smith

Despite growing attention to deepfake speech detection, the aspects of bias and fairness remain underexplored in the speech domain. To address this gap, we introduce the Speaker Characteristics Deepfake (SCDF) dataset: a novel, richly…

声音 · 计算机科学 2025-08-12 Vojtěch Staněk , Karel Srna , Anton Firc , Kamil Malinka

In this work, we open up the DAWT dataset - Densely Annotated Wikipedia Texts across multiple languages. The annotations include labeled text mentions mapping to entities (represented by their Freebase machine ids) as well as the type of…

信息检索 · 计算机科学 2017-03-06 Nemanja Spasojevic , Preeti Bhargava , Guoning Hu
‹ 上一页 1 8 9 10 下一页 ›