English
Related papers

Related papers: Generating Synthetic Oracle Datasets to Analyze No…

200 papers

Smart sensing provides an easier and convenient data-driven mechanism for monitoring and control in the built environment. Data generated in the built environment are privacy sensitive and limited. Federated learning is an emerging paradigm…

Machine Learning · Computer Science 2022-09-07 Rahul Mishra , Hari Prabhat Gupta , Tanima Dutta , Sajal K. Das

In this paper, we investigate the issue of detecting the real-life influence of people based on their Twitter account. We propose an overview of common Twitter features used to characterize such accounts and their activity, and show that…

Social and Information Networks · Computer Science 2021-08-06 Jean-Val{è}re Cossu , Nicolas Dugu{é} , Vincent Labatut

Applying natural language processing for mining and intelligent information access to tweets (a form of microblog) is a challenging, emerging research area. Unlike carefully authored news text and other longer content, tweets pose a number…

Computation and Language · Computer Science 2014-11-26 Leon Derczynski , Diana Maynard , Giuseppe Rizzo , Marieke van Erp , Genevieve Gorrell , Raphaël Troncy , Johann Petrak , Kalina Bontcheva

Personal attacks in the context of social media conversations often lead to fast-paced derailment, leading to even more harmful exchanges being made. State-of-the-art systems for the detection of such conversational derailment often make…

Computation and Language · Computer Science 2023-11-20 Steven Leung , Filippos Papapolyzos

In the last couple decades, social network services like Twitter have generated large volumes of data about users and their interests, providing meaningful business intelligence so organizations can better understand and engage their…

Computation and Language · Computer Science 2017-12-01 Angela Lin

Social media is often utilized as a lifeline for communication during natural disasters. Traditionally, natural disaster tweets are filtered from the Twitter stream using the name of the natural disaster and the filtered tweets are sent for…

Computation and Language · Computer Science 2022-07-12 Ramya Tekumalla , Juan M. Banda

Improving Sparse Autoencoders (SAEs) requires benchmarks that can precisely validate architectural innovations. However, current SAE benchmarks on LLMs are often too noisy to differentiate architectural improvements, and current synthetic…

Machine Learning · Computer Science 2026-02-17 David Chanin , Adrià Garriga-Alonso

Crowdsourcing platforms are often used to collect datasets for training machine learning models, despite higher levels of inaccurate labeling compared to expert labeling. There are two common strategies to manage the impact of such noise.…

Computation and Language · Computer Science 2022-06-14 Derek Chen , Zhou Yu , Samuel R. Bowman

A comprehensive understanding of data quality is the cornerstone of measurement studies in social media research. This paper presents in-depth measurements on the effects of Twitter data sampling across different timescales and different…

Social and Information Networks · Computer Science 2020-04-07 Siqi Wu , Marian-Andrei Rizoiu , Lexing Xie

Large Language Models (LLMs) have demonstrated impressive performance across various tasks, including sentiment analysis. However, data quality--particularly when sourced from social media--can significantly impact their accuracy. This…

Computation and Language · Computer Science 2025-04-09 Naman Bhargava , Mohammed I. Radaideh , O Hwang Kwon , Aditi Verma , Majdi I. Radaideh

Tweet clustering for event detection is a powerful modern method to automate the real-time detection of events. In this work we present a new tweet clustering approach, using a probabilistic approach to incorporate temporal information. By…

Social and Information Networks · Computer Science 2018-11-14 Peter Mathews , Caitlin Gray , Lewis Mitchell , Giang T. Nguyen , Nigel G. Bean

The development of data-driven heart sound classification models has been an active area of research in recent years. To develop such data-driven models in the first place, heart sound signals need to be captured using a signal acquisition…

Audio and Speech Processing · Electrical Eng. & Systems 2022-11-15 Davoud Shariat Panah , Andrew Hines , Susan McKeever

In this paper we describe our attempt at producing a state-of-the-art Twitter sentiment classifier using Convolutional Neural Networks (CNNs) and Long Short Term Memory (LSTMs) networks. Our system leverages a large amount of unlabeled data…

Computation and Language · Computer Science 2017-04-21 Mathieu Cliche

Crowd counting is a critical task in computer vision, with several important applications. However, existing counting methods rely on labor-intensive density map annotations, necessitating the manual localization of each individual…

Computer Vision and Pattern Recognition · Computer Science 2023-10-13 Adriano D'Alessandro , Ali Mahdavi-Amiri , Ghassan Hamarneh

Generalizability is the ultimate goal of Machine Learning (ML) image classifiers, for which noise and limited dataset size are among the major concerns. We tackle these challenges through utilizing the framework of deep Multitask Learning…

Computer Vision and Pattern Recognition · Computer Science 2022-08-23 Khashayar Namdar , Partoo Vafaeikia , Farzad Khalvati

As malicious actors employ increasingly advanced and widespread bots to disseminate misinformation and manipulate public opinion, the detection of Twitter bots has become a crucial task. Though graph-based Twitter bot detection methods…

Artificial Intelligence · Computer Science 2024-01-04 Zijian Cai , Zhaoxuan Tan , Zhenyu Lei , Zifeng Zhu , Hongrui Wang , Qinghua Zheng , Minnan Luo

Traditional semantic similarity models often fail to encapsulate the external context in which texts are situated. However, textual datasets generated on mobile platforms can help us build a truer representation of semantic similarity by…

Computation and Language · Computer Science 2018-12-27 Peter Hansel , Nik Marda , William Yin

Extracting noisy or incorrectly labeled samples from a labeled dataset with hard/difficult samples is an important yet under-explored topic. Two general and often independent lines of work exist, one focuses on addressing noisy labels, and…

Machine Learning · Computer Science 2023-07-21 Mahsa Forouzesh , Patrick Thiran

Recent advances in large language models (LLMs) have demonstrated strong performance on simple text classification tasks, frequently under zero-shot settings. However, their efficacy declines when tackling complex social media challenges…

Computation and Language · Computer Science 2025-04-23 Elyas Meguellati , Assaad Zeghina , Shazia Sadiq , Gianluca Demartini

Robustness to label noise is a critical property for weakly-supervised classifiers trained on massive datasets. Robustness to label noise is a critical property for weakly-supervised classifiers trained on massive datasets. In this paper,…

Machine Learning · Computer Science 2020-07-14 Amirmasoud Ghiassi , Taraneh Younesian , Robert Birke , Lydia Y. Chen