中文
相关论文

相关论文: Small data problems in political research: a criti…

200 篇论文

A large proportion of online comments present on public domains are constructive, however a significant proportion are toxic in nature. The comments contain lot of typos which increases the number of features manifold, making the ML model…

计算与语言 · 计算机科学 2018-08-31 Fahim Mohammad

Mapping political party systems to metric policy spaces is one of the major methodological problems in political science. At present, in most political science project this task is performed by domain experts relying on purely qualitative…

Sentiment Analysis of microblog feeds has attracted considerable interest in recent times. Most of the current work focuses on tweet sentiment classification. But not much work has been done to explore how reliable the opinions of the mass…

机器学习 · 计算机科学 2019-12-12 Rahul Radhakrishnan Iyer , Ronghuo Zheng , Yuezhang Li , Katia Sycara

Sentiment classification in short text datasets faces significant challenges such as class imbalance, limited training samples, and the inherent subjectivity of sentiment labels -- issues that are further intensified by the limited context…

计算与语言 · 计算机科学 2025-09-08 Julius Neumann , Robert Lange , Yuni Susanti , Michael Färber

There are some shreds of evidence that social opinion polarization leads to the breakup of the relationship, some in the scale of small communities, but others can divide large organizations or even a nation. The legacy methodology to…

社会与信息网络 · 计算机科学 2021-02-17 Andry Alamsyah , Wachda Yuniar Rochmah , Arina Nahya Nurnafia

Research shows that exposure to suicide-related news media content is associated with suicide rates, with some content characteristics likely having harmful and others potentially protective effects. Although good evidence exists for a few…

计算与语言 · 计算机科学 2022-06-29 Hannah Metzler , Hubert Baginski , Thomas Niederkrotenthaler , David Garcia

For text classification tasks, finetuned language models perform remarkably well. Yet, they tend to rely on spurious patterns in training data, thus limiting their performance on out-of-distribution (OOD) test data. Among recent models…

计算与语言 · 计算机科学 2022-10-24 Maarten De Raedt , Fréderic Godin , Chris Develder , Thomas Demeester

In recent studies [1][13][12] Recurrent Neural Networks were used for generative processes and their surprising performance can be explained by their ability to create good predictions. In addition, data compression is also based on…

计算与语言 · 计算机科学 2017-05-03 Juan Andrés Laura , Gabriel Masi , Luis Argerich

Well curated, large-scale corpora of social media posts containing broad public opinion offer an alternative data source to complement traditional surveys. While surveys are effective at collecting representative samples and are capable of…

计算与语言 · 计算机科学 2025-02-14 Michael V. Arnold , Peter Sheridan Dodds , Christopher M. Danforth

A public dataset, with a variety of properties suitable for sentiment analysis [1], event prediction, trend detection and other text mining applications, is needed in order to be able to successfully perform analysis studies. The vast…

计算与语言 · 计算机科学 2018-02-01 Semiha Makinist , Ibrahim Riza Hallac , Betul Ay Karakus , Galip Aydin

The role of social media in opinion formation has far-reaching implications in all spheres of society. Though social media provide platforms for expressing news and views, it is hard to control the quality of posts due to the sheer volumes…

机器学习 · 计算机科学 2021-09-08 Rini Anggrainingsih , Ghulam Mubashar Hassan , Amitava Datta

Answer sentence selection (AS2) modeling requires annotated data, i.e., hand-labeled question-answer pairs. We present a strategy to collect weakly supervised answers for a question based on its reference to improve AS2 modeling.…

计算与语言 · 计算机科学 2021-04-20 Vivek Krishnamurthy , Thuy Vu , Alessandro Moschitti

Big Data dealing with the social produce predictive correlations for the benefit of brands and web platforms. Beyond "society" and "opinion" for which the text lays out a genealogy, appear the "traces" that must be theorized as…

社会与信息网络 · 计算机科学 2016-07-19 Dominique Boullier

Twitter is currently one of the biggest social media platforms. Its users may share, read, and engage with short posts called tweets. For the ACM Recommender Systems Conference 2020, Twitter published a dataset around 70 GB in size for the…

信息检索 · 计算机科学 2023-10-06 Jovan Jeromela

Pre-trained language models have shown excellent results in few-shot learning scenarios using in-context learning. Although it is impressive, the size of language models can be prohibitive to make them usable in on-device applications, such…

计算与语言 · 计算机科学 2022-04-27 Navid Rezaei , Marek Z. Reformat

Large language models (LLMs) have gained significant attention due to their ability to mimic human language. Identifying texts generated by LLMs is crucial for understanding their capabilities and mitigating potential consequences. This…

计算与语言 · 计算机科学 2024-07-19 Anjali Rawal , Hui Wang , Youjia Zheng , Yu-Hsuan Lin , Shanu Sushmita

Microblogs such as Twitter represent a powerful source of information. Part of this information can be aggregated beyond the level of individual posts. Some of this aggregated information is referring to events that could or should be acted…

计算与语言 · 计算机科学 2020-08-04 Ali Hürriyetoğlu

As the number of applications that use machine learning algorithms increases, the need for labeled data useful for training such algorithms intensifies. Getting labels typically involves employing humans to do the annotation, which directly…

机器学习 · 计算机科学 2013-07-16 Alexandros Ntoulas , Omar Alonso , Vasilis Kandylas

To obtain a large amount of training labels inexpensively, researchers have recently adopted the weak supervision (WS) paradigm, which leverages labeling rules to synthesize training labels rather than using individual annotations to…

计算与语言 · 计算机科学 2022-10-10 Linxin Song , Jieyu Zhang , Tianxiang Yang , Masayuki Goto

The growing need to analyze large collections of documents has led to great developments in topic modeling. Since documents are frequently associated with other related variables, such as labels or ratings, much interest has been placed on…

机器学习 · 统计学 2018-08-20 Filipe Rodrigues , Mariana Lourenço , Bernardete Ribeiro , Francisco Pereira