English
Related papers

Related papers: ILiAD: An Interactive Corpus for Linguistic Annota…

200 papers

Inspiration moves a person to see new possibilities and transforms the way they perceive their own potential. Inspiration has received little attention in psychology, and has not been researched before in the NLP community. To the best of…

Computation and Language · Computer Science 2023-05-31 Oana Ignat , Y-Lan Boureau , Jane A. Yu , Alon Halevy

The use of irony and sarcasm in social media allows us to study them at scale for the first time. However, their diversity has made it difficult to construct a high-quality corpus of sarcasm in dialogue. Here, we describe the process of…

Computation and Language · Computer Science 2017-09-19 Shereen Oraby , Vrindavan Harrison , Lena Reed , Ernesto Hernandez , Ellen Riloff , Marilyn Walker

In this paper, we present a corpus for use in automatic readability assessment and automatic text simplification of German. The corpus is compiled from web sources and consists of approximately 211,000 sentences. As a novel contribution, it…

Computation and Language · Computer Science 2019-09-20 Alessia Battisti , Sarah Ebling

Twitter continuously tightens the access to its data via the publicly accessible, cost-free standard APIs. This especially applies to the follow network. In light of this, we successfully modified a network sampling method to work…

Social and Information Networks · Computer Science 2021-01-13 Felix Victor Münch , Ben Thies , Cornelius Puschmann , Axel Bruns

Modern society habitually uses online social media services to publicly share observations, thoughts, opinions, and beliefs at any time and from any location. These geotagged social media posts may provide aggregate insights into people's…

Social and Information Networks · Computer Science 2014-11-25 Derek Doran , Swapna Gokhale , Aldo Dagnino

Relevant and timely information collected from social media during crises can be an invaluable resource for emergency management. However, extracting this information remains a challenging task, particularly when dealing with social media…

Information Retrieval · Computer Science 2022-04-22 Fedor Vitiugin , Carlos Castillo

With the rapid expansion of content on social media platforms, analyzing and comprehending online discourse has become increasingly complex. This paper introduces LLMTaxo, a novel framework leveraging large language models for the automated…

Computation and Language · Computer Science 2025-10-21 Haiqi Zhang , Zhengyuan Zhu , Zeyu Zhang , Chengkai Li

Emergency-relevant data comes in many varieties. It can be high volume and high velocity, and reaction times are critical, calling for efficient and powerful techniques for data analysis and management. Knowledge graphs represent data in a…

Computers and Society · Computer Science 2021-01-18 Andreas L Opdahl

Information about pretraining corpora used to train the current best-performing language models is seldom discussed: commercial models rarely detail their data, and even open models are often released without accompanying training data or…

Identifying cohorts of patients based on eligibility criteria such as medical conditions, procedures, and medication use is critical to recruitment for clinical trials. Such criteria are often most naturally described in free-text, using…

Computation and Language · Computer Science 2022-07-29 Nicholas J Dobbins , Tony Mullen , Ozlem Uzuner , Meliha Yetisgen

Computer-mediated communication is driving fundamental changes in the nature of written language. We investigate these changes by statistical analysis of a dataset comprising 107 million Twitter messages (authored by 2.7 million unique user…

Computation and Language · Computer Science 2014-11-25 Jacob Eisenstein , Brendan O'Connor , Noah A. Smith , Eric P. Xing

A word embedding is a low-dimensional, dense and real- valued vector representation of a word. Word embeddings have been used in many NLP tasks. They are usually gener- ated from a large text corpus. The embedding of a word cap- tures both…

Computation and Language · Computer Science 2017-08-15 Quanzhi Li , Sameena Shah , Xiaomo Liu , Armineh Nourbakhsh

Social media is daily creating massive multimedia content with paired image and text, presenting the pressing need to automate the vision and language understanding for various multimodal classification tasks. Compared to the commonly…

Computation and Language · Computer Science 2023-03-28 Chunpu Xu , Jing Li

The increasing availability of digital collections of historical and contemporary literature presents a wealth of possibilities for new research in the humanities. The scale and diversity of such collections however, presents particular…

Computation and Language · Computer Science 2023-06-16 Susan Leavy , Gerardine Meaney , Karen Wade , Derek Greene

There is an ever growing number of users with accounts on multiple social media and networking sites. Consequently, there is increasing interest in matching user accounts and profiles across different social networks in order to create…

Social and Information Networks · Computer Science 2016-06-21 Soroush Vosoughi , Helen Zhou , Deb Roy

The increasing capacities of large language models (LLMs) have been shown to present an unprecedented opportunity to scale up data analytics in the humanities and social sciences, by automating complex qualitative tasks otherwise typically…

Computation and Language · Computer Science 2024-10-22 Andres Karjus

Introduction: Part-of-Speech (POS) Tagging, the process of classifying words into their respective parts of speech (e.g., verb or noun), is essential in various natural language processing applications. POS tagging is a crucial…

Computation and Language · Computer Science 2023-10-03 Leyla Rabiei , Farzaneh Rahmani , Mohammad Khansari , Zeinab Rajabi , Moein Salimi

Language Identification (LID) is a challenging task, especially when the input texts are short and noisy such as posts and statuses on social media or chat logs on gaming forums. The task has been tackled by either designing a feature set…

Computation and Language · Computer Science 2019-10-16 Duy Tin Vo , Richard Khoury

When speaking or writing, people omit information that seems clear and evident, such that only part of the message is expressed in words. Especially in argumentative texts it is very common that (important) parts of the argument are implied…

Computation and Language · Computer Science 2019-12-24 Maria Becker , Katharina Korfhage , Anette Frank

We present M2D2, a fine-grained, massively multi-domain corpus for studying domain adaptation in language models (LMs). M2D2 consists of 8.5B tokens and spans 145 domains extracted from Wikipedia and Semantic Scholar. Using ontologies…

Computation and Language · Computer Science 2022-10-17 Machel Reid , Victor Zhong , Suchin Gururangan , Luke Zettlemoyer
‹ Prev 1 8 9 10 Next ›