English
Related papers

Related papers: Bridging the Data Provenance Gap Across Text, Spee…

200 papers

To address the global challenge of online hate speech, prior research has developed detection models to flag such content on social media. However, due to systematic biases in evaluation datasets, the real-world effectiveness of these…

Computation and Language · Computer Science 2025-06-04 Manuel Tonneau , Diyi Liu , Niyati Malhotra , Scott A. Hale , Samuel P. Fraiberger , Victor Orozco-Olvera , Paul Röttger

The impressive capabilities of recent language models can be largely attributed to the multi-trillion token pretraining datasets that they are trained on. However, model developers fail to disclose their construction methodology which has…

The rapid development of Artificial Intelligence (AI) has revolutionized numerous fields, with large language models (LLMs) and computer vision (CV) systems driving advancements in natural language understanding and visual processing,…

Computation and Language · Computer Science 2024-12-04 Yunkai Dang , Kaichen Huang , Jiahao Huo , Yibo Yan , Sirui Huang , Dongrui Liu , Mengxi Gao , Jie Zhang , Chen Qian , Kun Wang , Yong Liu , Jing Shao , Hui Xiong , Xuming Hu

Misinformation is becoming increasingly prevalent on social media and in news articles. It has become so widespread that we require algorithmic assistance utilising machine learning to detect such content. Training these machine learning…

Machine Learning · Computer Science 2022-03-09 Dan Saattrup Nielsen , Ryan McConville

Despite the growing popularity of audio platforms, fact-checking spoken content remains significantly underdeveloped. Misinformation in speech often unfolds across multi-turn dialogues, shaped by speaker interactions, disfluencies,…

Social and Information Networks · Computer Science 2025-08-19 Chaewan Chun , Lysandre Terrisse , Delvin Ce Zhang , Dongwon Lee

The increasing use of machine learning models has amplified the demand for high-quality, large-scale multimodal datasets. However, the availability of such datasets, especially those combining acoustic, visual and textual data, remains…

Multimedia · Computer Science 2025-09-09 Jorge E. León , Miguel Carrasco

Automatic speech recognition (ASR) systems are designed to transcribe spoken language into written text and find utility in a variety of applications including voice assistants and transcription services. However, it has been observed that…

Computation and Language · Computer Science 2023-07-21 Anand Kumar Rai , Siddharth D Jaiswal , Animesh Mukherjee

Recent integration of Natural Language Processing (NLP) and multimodal models has advanced the field of sports analytics. This survey presents a comprehensive review of the datasets and applications driving these innovations post-2020. We…

Computation and Language · Computer Science 2024-06-19 Haotian Xia , Zhengbang Yang , Yun Zhao , Yuqing Wang , Jingxi Li , Rhys Tracy , Zhuangdi Zhu , Yuan-fang Wang , Hanjie Chen , Weining Shen

An increasing number of datasets contain multiple views, such as video, sound and automatic captions. A basic challenge in representation learning is how to leverage multiple views to learn better representations. This is further…

Machine Learning · Computer Science 2019-03-04 Nils Holzenberger , Shruti Palaskar , Pranava Madhyastha , Florian Metze , Raman Arora

We have now entered the era of trillion parameter machine learning models trained on billion-sized datasets scraped from the internet. The rise of these gargantuan datasets has given rise to formidable bodies of critical work that has…

Computers and Society · Computer Science 2021-10-06 Abeba Birhane , Vinay Uday Prabhu , Emmanuel Kahembwe

The rapid advancement of large language models has fundamentally shifted the bottleneck in AI development from computational power to data availability-with countless valuable datasets remaining hidden across specialized repositories,…

Artificial Intelligence · Computer Science 2025-08-12 Keyu Li , Mohan Jiang , Dayuan Fu , Yunze Wu , Xiangkun Hu , Dequan Wang , Pengfei Liu

It is well known that the quality and quantity of training data are significant factors which affect the development and performance of machine intelligence algorithms. Without representative data, neither scientists nor algorithms would be…

Machine Learning · Computer Science 2019-01-08 Georgios Mastorakis

The exponential growth of user-generated content on social media platforms has precipitated significant challenges in information management, particularly in content organization, retrieval, and discovery. Hashtags, as a fundamental…

Information Retrieval · Computer Science 2025-03-26 Shubhi Bansal , Kushaan Gowda , Anupama Sureshbabu K , Chirag Kothari , Nagendra Kumar

Pretraining from unlabelled web videos has quickly become the de-facto means of achieving high performance on many video understanding tasks. Features are learned via prediction of grounded relationships between visual content and automatic…

Computation and Language · Computer Science 2020-10-19 Jack Hessel , Zhenhai Zhu , Bo Pang , Radu Soricut

In recent years, discussions about fairness in machine learning, AI ethics and algorithm audits have increased. Many entities have developed framework guidance to establish a baseline rubric for fairness and accountability. However, in…

Machine Learning · Computer Science 2022-06-23 Cherie M Poland

With the advent of large multimodal language models, science is now at a threshold of an AI-based technological transformation. An emerging ecosystem of models and tools aims to support researchers throughout the scientific lifecycle,…

The rapid advancement of general-purpose AI models has increased concerns about copyright infringement in training data, yet current regulatory frameworks remain predominantly reactive rather than proactive. This paper examines the…

Computers and Society · Computer Science 2026-01-21 Mariia Kyrychenko , Mykyta Mudryi , Markiyan Chaklosh

We address the problem of predicting the leading political ideology, i.e., left-center-right bias, for YouTube channels of news media. Previous work on the problem has focused exclusively on text and on analysis of the language used, topics…

Computation and Language · Computer Science 2019-10-22 Yoan Dinkov , Ahmed Ali , Ivan Koychev , Preslav Nakov

Voice based technologies have the potential to bridge digital accessibility gaps; however, existing datasets fail to capture the linguistic and regional diversity of Indic languages. We present Project VAANI, a large scale multimodal…

‹ Prev 1 4 5 6 7 8 10 Next ›