中文
相关论文

相关论文: Multimodal datasets: misogyny, pornography, and ma…

200 篇论文

Text-conditioned image generation has gained significant attention in recent years and are processing increasingly longer and comprehensive text prompt. In everyday life, dense and intricate text appears in contexts like advertisements,…

计算机视觉与模式识别 · 计算机科学 2025-11-04 Alex Jinpeng Wang , Dongxing Mao , Jiawei Zhang , Weiming Han , Zhuobai Dong , Linjie Li , Yiqi Lin , Zhengyuan Yang , Libo Qin , Fuwei Zhang , Lijuan Wang , Min Li

Large Language Models (LLMs) are increasingly used to generate narrative content, including children's stories, which play an important role in social and cultural learning. Despite growing interest in AI safety and alignment, most existing…

计算与语言 · 计算机科学 2026-04-21 Yuxuan Ouyang , yingfeng luo , JingBo Zhu , Tong Xiao

Image generation has shown remarkable results in generating high-fidelity realistic images, in particular with the advancement of diffusion-based models. However, the prevalence of AI-generated images may have side effects for the machine…

计算机视觉与模式识别 · 计算机科学 2024-11-22 Maorong Wang , Nicolas Michel , Jiafeng Mao , Toshihiko Yamasaki

Machine learning models that convert user-written text descriptions into images are now widely available online and used by millions of users to generate millions of images a day. We investigate the potential for these models to amplify…

A recent study has shown that large-scale visual datasets are very biased: they can be easily classified by modern neural networks. However, the concrete forms of bias among these datasets remain unclear. In this study, we propose a…

计算机视觉与模式识别 · 计算机科学 2024-12-04 Boya Zeng , Yida Yin , Zhuang Liu

As demand for large corpora increases with the size of current state-of-the-art language models, using web data as the main part of the pre-training corpus for these models has become a ubiquitous practice. This, in turn, has introduced an…

计算与语言 · 计算机科学 2022-12-21 Tim Jansen , Yangling Tong , Victoria Zevallos , Pedro Ortiz Suarez

The detection of sensitive content in large datasets is crucial for ensuring that shared and analysed data is free from harmful material. However, current moderation tools, such as external APIs, suffer from limitations in customisation,…

计算与语言 · 计算机科学 2025-06-25 Dimosthenis Antypas , Indira Sen , Carla Perez-Almendros , Jose Camacho-Collados , Francesco Barbieri

One of the key challenges of detecting AI-generated images is spotting images that have been created by previously unseen generative models. We argue that the limited diversity of the training data is a major obstacle to addressing this…

计算机视觉与模式识别 · 计算机科学 2025-06-11 Jeongsoo Park , Andrew Owens

High-quality textual training data is essential for the success of multimodal data processing tasks, yet outputs from image captioning models like BLIP and GIT often contain errors and anomalies that are difficult to rectify using…

计算与语言 · 计算机科学 2025-02-25 Elyas Meguellati , Nardiena Pratama , Shazia Sadiq , Gianluca Demartini

In order to study online hate speech, the availability of datasets containing the linguistic phenomena of interest are of crucial importance. However, when it comes to specific target groups, for example teenagers, collecting such data may…

计算与语言 · 计算机科学 2020-05-06 Alessio Palmero Aprosio , Stefano Menini , Sara Tonelli

Large language models (LLMs) acquire general linguistic knowledge from massive-scale pretraining. However, pretraining data mainly comprised of web-crawled texts contain undesirable social biases which can be perpetuated or even amplified…

计算与语言 · 计算机科学 2025-09-04 Takuma Udagawa , Yang Zhao , Hiroshi Kanayama , Bishwaranjan Bhattacharjee

Online discussions, panels, talk page edits, etc., often contain harmful conversational content i.e., hate speech, death threats and offensive language, especially towards certain demographic groups. For example, individuals who identify as…

计算与语言 · 计算机科学 2022-07-21 Jamell Dacon , Harry Shomer , Shaylynn Crum-Dacon , Jiliang Tang

The last two years have seen a rapid growth in concerns around the safety of large language models (LLMs). Researchers and practitioners have met these concerns by creating an abundance of datasets for evaluating and improving LLM safety.…

计算与语言 · 计算机科学 2025-01-13 Paul Röttger , Fabio Pernisi , Bertie Vidgen , Dirk Hovy

With the advent of Large Language Models (LLMs) possessing increasingly impressive capabilities, a number of Large Vision-Language Models (LVLMs) have been proposed to augment LLMs with visual inputs. Such models condition generated text on…

计算机视觉与模式识别 · 计算机科学 2025-05-01 Phillip Howard , Kathleen C. Fraser , Anahita Bhiwandiwalla , Svetlana Kiritchenko

This study investigates ChatGPT-4o's multimodal content generation, highlighting significant disparities in its treatment of sexual content and nudity versus violent and drug-related themes. Detailed analysis reveals that ChatGPT-4o…

计算机与社会 · 计算机科学 2024-12-02 Roberto Balestri

In order to train, test, and evaluate nudity detection models, machine learning researchers typically rely on nude images scraped from the Internet. Our research finds that this content is collected and, in some cases, subsequently…

计算机与社会 · 计算机科学 2025-10-28 Princessa Cintaqia , Arshia Arya , Elissa M Redmiles , Deepak Kumar , Allison McDonald , Lucy Qin

Large language models (LLMs) rely heavily on web-scale datasets like Common Crawl, which provides over 80\% of training data for some modern models. However, the indiscriminate nature of web crawling raises challenges in data quality,…

计算与语言 · 计算机科学 2025-09-01 Inés Altemir Marinas , Anastasiia Kucherenko , Andrei Kucharavy

Large language models are commonly trained on a mixture of filtered web data and curated high-quality corpora, such as social media conversations, books, or technical papers. This curation process is believed to be necessary to produce…

Multi-modal large language models (MLLMs) have made significant progress, yet their safety alignment remains limited. Typically, current open-source MLLMs rely on the alignment inherited from their language module to avoid harmful…

密码学与安全 · 计算机科学 2025-04-15 Yanbo Wang , Jiyang Guan , Jian Liang , Ran He

Recent studies have demonstrated that large language models (LLMs) have ethical-related problems such as social biases, lack of moral reasoning, and generation of offensive content. The existing evaluation metrics and methods to address…

计算与语言 · 计算机科学 2024-02-23 Masahiro Kaneko , Danushka Bollegala , Timothy Baldwin