中文
相关论文

相关论文: Generating Synthetic Oracle Datasets to Analyze No…

200 篇论文

Smart sensing provides an easier and convenient data-driven mechanism for monitoring and control in the built environment. Data generated in the built environment are privacy sensitive and limited. Federated learning is an emerging paradigm…

机器学习 · 计算机科学 2022-09-07 Rahul Mishra , Hari Prabhat Gupta , Tanima Dutta , Sajal K. Das

In this paper, we investigate the issue of detecting the real-life influence of people based on their Twitter account. We propose an overview of common Twitter features used to characterize such accounts and their activity, and show that…

社会与信息网络 · 计算机科学 2021-08-06 Jean-Val{è}re Cossu , Nicolas Dugu{é} , Vincent Labatut

Applying natural language processing for mining and intelligent information access to tweets (a form of microblog) is a challenging, emerging research area. Unlike carefully authored news text and other longer content, tweets pose a number…

Personal attacks in the context of social media conversations often lead to fast-paced derailment, leading to even more harmful exchanges being made. State-of-the-art systems for the detection of such conversational derailment often make…

计算与语言 · 计算机科学 2023-11-20 Steven Leung , Filippos Papapolyzos

In the last couple decades, social network services like Twitter have generated large volumes of data about users and their interests, providing meaningful business intelligence so organizations can better understand and engage their…

计算与语言 · 计算机科学 2017-12-01 Angela Lin

Social media is often utilized as a lifeline for communication during natural disasters. Traditionally, natural disaster tweets are filtered from the Twitter stream using the name of the natural disaster and the filtered tweets are sent for…

计算与语言 · 计算机科学 2022-07-12 Ramya Tekumalla , Juan M. Banda

Improving Sparse Autoencoders (SAEs) requires benchmarks that can precisely validate architectural innovations. However, current SAE benchmarks on LLMs are often too noisy to differentiate architectural improvements, and current synthetic…

机器学习 · 计算机科学 2026-02-17 David Chanin , Adrià Garriga-Alonso

Crowdsourcing platforms are often used to collect datasets for training machine learning models, despite higher levels of inaccurate labeling compared to expert labeling. There are two common strategies to manage the impact of such noise.…

计算与语言 · 计算机科学 2022-06-14 Derek Chen , Zhou Yu , Samuel R. Bowman

A comprehensive understanding of data quality is the cornerstone of measurement studies in social media research. This paper presents in-depth measurements on the effects of Twitter data sampling across different timescales and different…

社会与信息网络 · 计算机科学 2020-04-07 Siqi Wu , Marian-Andrei Rizoiu , Lexing Xie

Large Language Models (LLMs) have demonstrated impressive performance across various tasks, including sentiment analysis. However, data quality--particularly when sourced from social media--can significantly impact their accuracy. This…

计算与语言 · 计算机科学 2025-04-09 Naman Bhargava , Mohammed I. Radaideh , O Hwang Kwon , Aditi Verma , Majdi I. Radaideh

Tweet clustering for event detection is a powerful modern method to automate the real-time detection of events. In this work we present a new tweet clustering approach, using a probabilistic approach to incorporate temporal information. By…

社会与信息网络 · 计算机科学 2018-11-14 Peter Mathews , Caitlin Gray , Lewis Mitchell , Giang T. Nguyen , Nigel G. Bean

The development of data-driven heart sound classification models has been an active area of research in recent years. To develop such data-driven models in the first place, heart sound signals need to be captured using a signal acquisition…

音频与语音处理 · 电气工程与系统科学 2022-11-15 Davoud Shariat Panah , Andrew Hines , Susan McKeever

In this paper we describe our attempt at producing a state-of-the-art Twitter sentiment classifier using Convolutional Neural Networks (CNNs) and Long Short Term Memory (LSTMs) networks. Our system leverages a large amount of unlabeled data…

计算与语言 · 计算机科学 2017-04-21 Mathieu Cliche

Crowd counting is a critical task in computer vision, with several important applications. However, existing counting methods rely on labor-intensive density map annotations, necessitating the manual localization of each individual…

计算机视觉与模式识别 · 计算机科学 2023-10-13 Adriano D'Alessandro , Ali Mahdavi-Amiri , Ghassan Hamarneh

Generalizability is the ultimate goal of Machine Learning (ML) image classifiers, for which noise and limited dataset size are among the major concerns. We tackle these challenges through utilizing the framework of deep Multitask Learning…

计算机视觉与模式识别 · 计算机科学 2022-08-23 Khashayar Namdar , Partoo Vafaeikia , Farzad Khalvati

As malicious actors employ increasingly advanced and widespread bots to disseminate misinformation and manipulate public opinion, the detection of Twitter bots has become a crucial task. Though graph-based Twitter bot detection methods…

人工智能 · 计算机科学 2024-01-04 Zijian Cai , Zhaoxuan Tan , Zhenyu Lei , Zifeng Zhu , Hongrui Wang , Qinghua Zheng , Minnan Luo

Traditional semantic similarity models often fail to encapsulate the external context in which texts are situated. However, textual datasets generated on mobile platforms can help us build a truer representation of semantic similarity by…

计算与语言 · 计算机科学 2018-12-27 Peter Hansel , Nik Marda , William Yin

Extracting noisy or incorrectly labeled samples from a labeled dataset with hard/difficult samples is an important yet under-explored topic. Two general and often independent lines of work exist, one focuses on addressing noisy labels, and…

机器学习 · 计算机科学 2023-07-21 Mahsa Forouzesh , Patrick Thiran

Recent advances in large language models (LLMs) have demonstrated strong performance on simple text classification tasks, frequently under zero-shot settings. However, their efficacy declines when tackling complex social media challenges…

计算与语言 · 计算机科学 2025-04-23 Elyas Meguellati , Assaad Zeghina , Shazia Sadiq , Gianluca Demartini

Robustness to label noise is a critical property for weakly-supervised classifiers trained on massive datasets. Robustness to label noise is a critical property for weakly-supervised classifiers trained on massive datasets. In this paper,…

机器学习 · 计算机科学 2020-07-14 Amirmasoud Ghiassi , Taraneh Younesian , Robert Birke , Lydia Y. Chen