中文
相关论文

相关论文: Multimodal datasets: misogyny, pornography, and ma…

200 篇论文

The advent of Large Language Models (LLMs) and generative AI is fundamentally transforming information retrieval and processing on the Internet, bringing both great potential and significant concerns regarding content authenticity and…

State-of-the-art face recognition models show impressive accuracy, achieving over 99.8% on Labeled Faces in the Wild (LFW) dataset. Such models are trained on large-scale datasets that contain millions of real human face images collected…

计算机视觉与模式识别 · 计算机科学 2022-10-07 Gwangbin Bae , Martin de La Gorce , Tadas Baltrusaitis , Charlie Hewitt , Dong Chen , Julien Valentin , Roberto Cipolla , Jingjing Shen

Contrastive pre-training on large-scale image-text pair datasets has driven major advances in vision-language representation learning. Recent work shows that pretraining on global data followed by language or culture specific fine-tuning is…

计算机视觉与模式识别 · 计算机科学 2026-04-02 Issa Sugiura , Shuhei Kurita , Yusuke Oda , Daisuke Kawahara , Yasuo Okabe , Naoaki Okazaki

This paper presents WanJuan-CC, a safe and high-quality open-sourced English webtext dataset derived from Common Crawl data. The study addresses the challenges of constructing large-scale pre-training datasets for language models, which…

We investigate the impact of deep generative models on potential social biases in upcoming computer vision models. As the internet witnesses an increasing influx of AI-generated images, concerns arise regarding inherent biases that may…

计算机视觉与模式识别 · 计算机科学 2024-04-05 Tianwei Chen , Yusuke Hirota , Mayu Otani , Noa Garcia , Yuta Nakashima

Since Multimodal Large Language Models (MLLMs) are increasingly being integrated into everyday tools and intelligent agents, growing concerns have arisen regarding their possible output of unsafe contents, ranging from toxic language and…

机器学习 · 计算机科学 2026-04-08 Yuping Yan , Yuhan Xie , Yuanshuai Li , Yingchao Yu , Lingjuan Lyu , Yaochu Jin

The remarkable ease of use of diffusion models for image generation has led to a proliferation of synthetic content online. While these models are often employed for legitimate purposes, they are also used to generate fake images that…

计算机视觉与模式识别 · 计算机科学 2025-05-19 Giulia Bertazzini , Daniele Baracchi , Dasara Shullani , Isao Echizen , Alessandro Piva

Recent works have found evidence of gender bias in models of machine translation and coreference resolution using mostly synthetic diagnostic datasets. While these quantify bias in a controlled experiment, they often do so on a small scale…

计算与语言 · 计算机科学 2021-09-13 Shahar Levy , Koren Lazar , Gabriel Stanovsky

Due to the data-driven nature of current face identity (FaceID) customization methods, all state-of-the-art models rely on large-scale datasets containing millions of high-quality text-image pairs for training. However, none of these…

计算机视觉与模式识别 · 计算机科学 2025-03-28 Shuhe Wang , Xiaoya Li , Jiwei Li , Guoyin Wang , Xiaofei Sun , Bob Zhu , Han Qiu , Mo Yu , Shengjie Shen , Tianwei Zhang , Eduard Hovy

Massive web-crawled image-text datasets lay the foundation for recent progress in multimodal learning. These datasets are designed with the goal of training a model to do well on standard computer vision benchmarks, many of which, however,…

计算机视觉与模式识别 · 计算机科学 2025-09-16 Thao Nguyen , Matthew Wallingford , Sebastin Santy , Wei-Chiu Ma , Sewoong Oh , Ludwig Schmidt , Pang Wei Koh , Ranjay Krishna

Image captioning has made substantial progress with huge supporting image collections sourced from the web. However, recent studies have pointed out that captioning datasets, such as COCO, contain gender bias found in web corpora. As a…

计算机视觉与模式识别 · 计算机科学 2021-04-22 Ruixiang Tang , Mengnan Du , Yuening Li , Zirui Liu , Na Zou , Xia Hu

The QUILT-1M dataset is the first openly available dataset containing images harvested from various online sources. While it provides a huge data variety, the image quality and composition is highly heterogeneous, impacting its utility for…

计算机视觉与模式识别 · 计算机科学 2024-04-12 Marc Aubreville , Jonathan Ganz , Jonas Ammeling , Christopher C. Kaltenecker , Christof A. Bertram

Pornographic content occurring in human-machine interaction dialogues can cause severe side effects for users in open-domain dialogue systems. However, research on detecting pornographic language within human-machine interaction dialogues…

计算与语言 · 计算机科学 2024-03-21 Huachuan Qiu , Shuai Zhang , Hongliang He , Anqi Li , Zhenzhong Lan

We study the effectiveness of data-balancing for mitigating biases in contrastive language-image pretraining (CLIP), identifying areas of strength and limitation. First, we reaffirm prior conclusions that CLIP models can inadvertently…

机器学习 · 计算机科学 2024-03-08 Ibrahim Alabdulmohsin , Xiao Wang , Andreas Steiner , Priya Goyal , Alexander D'Amour , Xiaohua Zhai

The dissemination of Large Language Models (LLMs), trained at scale, and endowed with powerful text-generating abilities, has made it easier for all to produce harmful, toxic, faked or forged content. In response, various proposals have…

计算与语言 · 计算机科学 2025-06-12 Matthieu Dubois , François Yvon , Pablo Piantanida

Image content safety has become a significant challenge with the rise of visual media on online platforms. Meanwhile, in the age of AI-generated content (AIGC), many image generation models are capable of producing harmful content, such as…

Large Language Models (LLMs) inherit explicit and implicit biases from their training datasets. Identifying and mitigating biases in LLMs is crucial to ensure fair outputs, as they can perpetuate harmful stereotypes and misinformation. This…

机器学习 · 计算机科学 2025-11-19 Fatima Kazi , Alex Young , Yash Inani , Setareh Rafatirad

This paper investigates the challenges associated with bias, toxicity, unreliability, and lack of robustness in large language models (LLMs) such as ChatGPT. It emphasizes that these issues primarily stem from the quality and diversity of…

计算机与社会 · 计算机科学 2024-10-21 Federico Torrielli

With the ongoing rapid adoption of Artificial Intelligence (AI)-based systems in high-stakes domains, ensuring the trustworthiness, safety, and observability of these systems has become crucial. It is essential to evaluate and monitor AI…

计算与语言 · 计算机科学 2024-07-19 Krishnaram Kenthapadi , Mehrnoosh Sameki , Ankur Taly