中文
相关论文

相关论文: Understanding State Preferences With Text As Data:…

200 篇论文

Out of nearly 70,000 bills introduced in the U.S. Congress from 2001 to 2015, only 2,513 were enacted. We developed a machine learning approach to forecasting the probability that any bill will become law. Starting in 2001 with the 107th…

计算与语言 · 计算机科学 2017-07-05 John J. Nay

We examine a large dialog corpus obtained from the conversation history of a single individual with 104 conversation partners. The corpus consists of half a million instant messages, across several messaging platforms. We focus our analyses…

计算与语言 · 计算机科学 2019-04-29 Charles Welch , Verónica Pérez-Rosas , Jonathan K. Kummerfeld , Rada Mihalcea

The widespread use of social media has led to a surge in popularity for automated methods of analyzing public opinion. Supervised methods are adept at text categorization, yet the dynamic nature of social media discussions poses a continual…

计算与语言 · 计算机科学 2025-01-28 Tunazzina Islam , Dan Goldwasser

We present the Knesset Corpus, a corpus of Hebrew parliamentary proceedings containing over 30 million sentences (over 384 million tokens) from all the (plenary and committee) protocols held in the Israeli parliament between 1998 and 2022.…

计算与语言 · 计算机科学 2025-06-02 Gili Goldin , Nick Howell , Noam Ordan , Ella Rabinovich , Shuly Wintner

In this study, we use recent stance detection methods to study the stance (for, against or neutral) of statements in official information booklets for voters. Our main goal is to answer the fundamental question: are topics to be voted on…

计算与语言 · 计算机科学 2023-06-16 Eric Egli , Noah Mamié , Eyal Liron Dolev , Mathias Müller

Group deliberation enables people to collaborate and solve problems, however, it is understudied due to a lack of resources. To this end, we introduce the first publicly available dataset containing collaborative conversations on solving a…

计算与语言 · 计算机科学 2023-04-18 Georgi Karadzhov , Tom Stafford , Andreas Vlachos

We present a dataset of 833k paragraphs extracted from CC-BY licensed scientific publications, classified into four categories: acknowledgments, data mentions, software/code mentions, and clinical trial mentions. The paragraphs are…

计算与语言 · 计算机科学 2025-10-28 Eric Jeangirard

This paper introduces a computational framework designed to delineate gender distribution biases in topics covered by French TV and radio news. We transcribe a dataset of 11.7k hours, broadcasted in 2023 on 21 French channels. A Large…

计算与语言 · 计算机科学 2024-07-22 Valentin Pelloin , Lena Dodson , Émile Chapuis , Nicolas Hervé , David Doukhan

We analyze publicly available US Supreme Court documents using automated stance detection. In the first phase of our work, we investigate the extent to which the Court's public-facing language is political. We propose and calculate two…

计算与语言 · 计算机科学 2022-11-22 Noah Bergam , Emily Allaway , Kathleen McKeown

With the increasing usage of the internet, more and more data is being digitized including parliamentary debates but they are in an unstructured format. There is a need to convert them into a structured format for linguistic analysis. Much…

计算与语言 · 计算机科学 2018-08-22 Sakala Venkata Krishna Rohit , Navjyoti Singh

In this work, we create a web application to highlight the output of NLP models trained to parse and label discourse segments in law text. Our system is built primarily with journalists and legal interpreters in mind, and we focus on…

计算与语言 · 计算机科学 2022-07-01 Alexander Spangher , Jonathan May

We describe a gold standard corpus of protest events that comprise of various local and international sources from various countries in English. The corpus contains document, sentence, and token level annotations. This corpus facilitates…

计算与语言 · 计算机科学 2020-08-04 Ali Hürriyetoğlu , Erdem Yörük , Deniz Yüret , Osman Mutlu , Çağrı Yoltar , Fırat Duruşan , Burak Gürel

A U.S. Senator from South Dakota donated documents that were accumulated during his service as a house representative and senator to be housed at the Bridges library at South Dakota State University. This project investigated the utility of…

信息检索 · 计算机科学 2019-04-30 Damon Bayer , Semhar Michael

Argument mining and stance detection are central to understanding how opinions are formed and contested in online discourse. However, most publicly available resources focus on mainstream platforms such as Twitter and Reddit, leaving…

计算与语言 · 计算机科学 2026-02-17 Fathima Ameen , Danielle Brown , Manusha Malgareddy , Amanul Haque

Automatic machine learning systems can inadvertently accentuate and perpetuate inappropriate human biases. Past work on examining inappropriate biases has largely focused on just individual systems. Further, there is no benchmark dataset…

计算与语言 · 计算机科学 2018-05-14 Svetlana Kiritchenko , Saif M. Mohammad

Insightful findings in political science often require researchers to analyze documents of a certain subject or type, yet these documents are usually contained in large corpora that do not distinguish between pertinent and non-pertinent…

计算与语言 · 计算机科学 2019-10-29 Shrey Desai , Barea Sinno , Alex Rosenfeld , Junyi Jessy Li

The scientific paper output of the United Nations University (UNU) was bibliometrically analysed. It was found that (i) a noticeable continous paper output starts in 1995, (ii) about 65% of the research papers have been published as…

数字图书馆 · 计算机科学 2020-05-13 Johannes Stegmann

Social surveys have been widely used as a method of obtaining public opinion. Sometimes it is more ideal to collect opinions by presenting questions in free-response formats than in multiple-choice formats. Despite their advantages,…

社会与信息网络 · 计算机科学 2020-07-09 Tatsuro Kawamoto , Takaaki Aoki

Citing opinions is a powerful yet understudied strategy in argumentation. For example, an environmental activist might say, "Leading scientists agree that global warming is a serious concern," framing a clause which affirms their own stance…

计算与语言 · 计算机科学 2021-01-19 Yiwei Luo , Dallas Card , Dan Jurafsky

This paper introduces a document grounded dataset for text conversations. We define "Document Grounded Conversations" as conversations that are about the contents of a specified document. In this dataset the specified documents were…

计算与语言 · 计算机科学 2018-09-21 Kangyan Zhou , Shrimai Prabhumoye , Alan W Black