中文
相关论文

相关论文: GeniL: A Multilingual Dataset on Generalizing Lang…

200 篇论文

While generative multilingual models are rapidly being deployed, their safety and fairness evaluations are largely limited to resources collected in English. This is especially problematic for evaluations targeting inherently socio-cultural…

计算与语言 · 计算机科学 2024-03-12 Mukul Bhutani , Kevin Robinson , Vinodkumar Prabhakaran , Shachi Dave , Sunipa Dev

Generic sentences express generalisations about the world without explicit quantification. Although generics are central to everyday communication, building a precise semantic framework has proven difficult, in part because speakers use…

计算与语言 · 计算机科学 2024-12-17 Gustavo Cilleruelo Calderón , Emily Allaway , Barry Haddow , Alexandra Birch

Large language models (LLMs) have been shown to propagate and amplify harmful stereotypes, particularly those that disproportionately affect marginalised communities. To understand the effect of these stereotypes more comprehensively, we…

计算与语言 · 计算机科学 2024-10-10 Zara Siddique , Liam D. Turner , Luis Espinosa-Anke

A stereotype is a generalized perception of a specific group of humans. It is often potentially encoded in human language, which is more common in texts on social issues. Previous works simply define a sentence as stereotypical and…

计算与语言 · 计算机科学 2024-01-30 Yang Liu

Hate speech detection has become an important research topic within the past decade. More private corporations are needing to regulate user generated content on different platforms across the globe. In this paper, we introduce a study of…

计算与语言 · 计算机科学 2022-01-28 Neha Deshpande , Nicholas Farris , Vidhur Kumar

MGen is a dataset of over 4 million naturally occurring generic and quantified sentences extracted from diverse textual sources. Sentences in the dataset have long context documents, corresponding to websites and academic papers, and cover…

计算与语言 · 计算机科学 2025-11-25 Gustavo Cilleruelo , Emily Allaway , Barry Haddow , Alexandra Birch

Generative AI, such as large language models, has undergone rapid development within recent years. As these models become increasingly available to the public, concerns arise about perpetuating and amplifying harmful biases in applications.…

计算与语言 · 计算机科学 2024-09-04 Sara Sterlie , Nina Weng , Aasa Feragen

Natural language generation models are computer systems that generate coherent language when prompted with a sequence of words as context. Despite their ubiquity and many beneficial applications, language generation models also have the…

计算与语言 · 计算机科学 2022-07-25 Amin Rasekh , Ian Eisenberg

Multilingual studies of social bias in open-ended LLM generation remain limited: most existing benchmarks are English-centric, template-based, or restricted to recognizing pre-specified stereotypes. We introduce StereoTales, a multilingual…

Generic `toxicity' classifiers continue to be used for evaluating the potential for harm in natural language generation, despite mounting evidence of their shortcomings. We consider the challenge of measuring misogyny in natural language…

计算与语言 · 计算机科学 2023-12-07 Aaron J. Snoswell , Lucinda Nelson , Hao Xue , Flora D. Salim , Nicolas Suzor , Jean Burgess

AI-text detectors achieve high accuracy on in-domain benchmarks, but often struggle to generalize across different generation conditions such as unseen prompts, model families, or domains. While prior work has reported these generalization…

计算与语言 · 计算机科学 2026-01-27 Yuxi Xia , Kinga Stańczak , Benjamin Roth

Large Language Models (LLM) have made significant advances in the recent past becoming more mainstream in Artificial Intelligence (AI) enabled human-facing applications. However, LLMs often generate stereotypical output inherited from…

计算与语言 · 计算机科学 2023-11-27 Wu Zekun , Sahan Bulathwela , Adriano Soares Koshiyama

Many natural language inference (NLI) datasets contain biases that allow models to perform well by only using a biased subset of the input, without considering the remainder features. For instance, models are able to make a classification…

计算与语言 · 计算机科学 2021-09-01 Dimion Asael , Zachary Ziegler , Yonatan Belinkov

Social media often serves as a breeding ground for various hateful and offensive content. Identifying such content on social media is crucial due to its impact on the race, gender, or religion in an unprejudiced society. However, while…

计算与语言 · 计算机科学 2022-10-10 Mithun Das , Somnath Banerjee , Punyajoy Saha , Animesh Mukherjee

Stereotype benchmark datasets are crucial to detect and mitigate social stereotypes about groups of people in NLP models. However, existing datasets are limited in size and coverage, and are largely restricted to stereotypes prevalent in…

计算与语言 · 计算机科学 2023-05-22 Akshita Jha , Aida Davani , Chandan K. Reddy , Shachi Dave , Vinodkumar Prabhakaran , Sunipa Dev

Recently, commonsense reasoning in text generation has attracted much attention. Generative commonsense reasoning is the task that requires machines, given a group of keywords, to compose a single coherent sentence with commonsense…

计算与语言 · 计算机科学 2023-10-31 Yunxiang Zhang , Xiaojun Wan

Progress in natural language generation research has been shaped by the ever-growing size of language models. While large language models pre-trained on web data can generate human-sounding text, they also reproduce social biases and…

计算与语言 · 计算机科学 2023-06-06 Celine Wald , Lukas Pfahler

The surge in popularity of large language models has given rise to concerns about biases that these models could learn from humans. We investigate whether ingroup solidarity and outgroup hostility, fundamental social identity biases known…

计算与语言 · 计算机科学 2024-06-18 Tiancheng Hu , Yara Kyrychenko , Steve Rathje , Nigel Collier , Sander van der Linden , Jon Roozenbeek

Dehumanization, i.e., denying human qualities to individuals or groups, is a particularly harmful form of hate speech that can normalize violence against marginalized communities. Despite advances in NLP for detecting general hate speech,…

计算与语言 · 计算机科学 2025-07-11 Hamidreza Saffari , Mohammadamin Shafiei , Hezhao Zhang , Lasana Harris , Nafise Sadat Moosavi

The proliferation of online hate speech poses a significant threat to the harmony of the web. While explicit hate is easily recognized through overt slurs, implicit hate speech is often conveyed through sarcasm, irony, stereotypes, or coded…

计算与语言 · 计算机科学 2026-02-04 Chengshuai Zhao , Shu Wan , Paras Sheth , Karan Patwa , K. Selçuk Candan , Huan Liu
‹ 上一页 1 2 3 10 下一页 ›