English
Related papers

Related papers: BiaSWE: An Expert Annotated Dataset for Misogyny D…

200 papers

Algorithms are widely applied to detect hate speech and abusive language in social media. We investigated whether the human-annotated data used to train these algorithms are biased. We utilized a publicly available annotated Twitter dataset…

Computation and Language · Computer Science 2020-05-29 Jae Yeon Kim , Carlos Ortiz , Sarah Nam , Sarah Santiago , Vivek Datta

Current research on hate speech analysis is typically oriented towards monolingual and single classification tasks. In this paper, we present a new multilingual hate speech analysis dataset for English, Hindi, Arabic, French, German and…

Computation and Language · Computer Science 2023-04-04 Ankit Yadav , Shubham Chandel , Sushant Chatufale , Anil Bandhakavi

With the growth of online services, the need for advanced text classification algorithms, such as sentiment analysis and biased text detection, has become increasingly evident. The anonymous nature of online services often leads to the…

Computation and Language · Computer Science 2023-11-14 Dasol Choi , Jooyoung Song , Eunsun Lee , Jinwoo Seo , Heejune Park , Dongbin Na

The growing importance of culturally-aware natural language processing systems has led to an increasing demand for resources that capture sociopragmatic phenomena across diverse languages. Nevertheless, Arabic-language resources for…

As the body of research on abusive language detection and analysis grows, there is a need for critical consideration of the relationships between different subtasks that have been grouped under this label. Based on work on hate speech,…

Computation and Language · Computer Science 2017-05-31 Zeerak Waseem , Thomas Davidson , Dana Warmsley , Ingmar Weber

Warning: this work contains upsetting or disturbing content. Large language models (LLMs) tend to learn the social and cultural biases present in the raw pre-training data. To test if an LLM's behavior is fair, functional datasets are…

Computation and Language · Computer Science 2024-03-27 Veronika Grigoreva , Anastasiia Ivanova , Ilseyar Alimova , Ekaterina Artemova

Artificial Intelligence (AI) software systems, such as Sentiment Analysis (SA) systems, typically learn from large amounts of data that may reflect human biases. Consequently, the machine learning model in such software systems may exhibit…

Software Engineering · Computer Science 2021-10-06 Muhammad Hilmi Asyrofi , Zhou Yang , Imam Nur Bani Yusuf , Hong Jin Kang , Ferdian Thung , David Lo

Annotated data is an essential ingredient in natural language processing for training and evaluating machine learning models. It is therefore very desirable for the annotations to be of high quality. Recent work, however, has shown that…

Computation and Language · Computer Science 2022-09-27 Jan-Christoph Klie , Bonnie Webber , Iryna Gurevych

Speech style editing refers to modifying the stylistic properties of speech while preserving its linguistic content and speaker identity. However, most existing approaches depend on explicit labels or reference audio, which limits both…

Audio and Speech Processing · Electrical Eng. & Systems 2025-09-30 Yun Chen , Qi Chen , Zheqi Dai , Arshdeep Singh , Philip J. B. Jackson , Mark D. Plumbley

This paper is concerned with automatic continuous speech recognition using trainable systems. The aim of this work is to build acoustic models for spoken Swedish. This is done employing hidden Markov models and using the SpeechDat database…

Audio and Speech Processing · Electrical Eng. & Systems 2024-04-26 Giampiero Salvi

While understanding and removing gender biases in language models has been a long-standing problem in Natural Language Processing, prior research work has primarily been limited to English. In this work, we investigate some of the…

Computation and Language · Computer Science 2023-07-06 Aniket Vashishtha , Kabir Ahuja , Sunayana Sitaram

Media coverage has a substantial effect on the public perception of events. Nevertheless, media outlets are often biased. One way to bias news articles is by altering the word choice. The automatic identification of bias by word choice is…

Computation and Language · Computer Science 2023-05-04 Timo Spinde , Manuel Plank , Jan-David Krieger , Terry Ruas , Bela Gipp , Akiko Aizawa

This work presents a suite of fine-tuned Whisper models for Swedish, trained on a dataset of unprecedented size and variability for this mid-resourced language. As languages of smaller sizes are often underrepresented in multilingual…

Computation and Language · Computer Science 2025-08-15 Leonora Vesterbacka , Faton Rekathati , Robin Kurtz , Justyna Sikora , Agnes Toftgård

Suicidal ideation detection is critical for real-time suicide prevention, yet its progress faces two under-explored challenges: limited language coverage and unreliable annotation practices. Most available datasets are in English, but even…

Computation and Language · Computer Science 2025-07-22 Amina Dzafic , Merve Kavut , Ulya Bayram

This paper presents work on detecting misogyny in the comments of a large Austrian German language newspaper forum. We describe the creation of a corpus of 6600 comments which were annotated with 5 levels of misogyny. The forum moderators…

Computation and Language · Computer Science 2022-12-01 Johann Petrak , Brigitte Krenn

High annotation costs from hiring or crowdsourcing complicate the creation of large, high-quality datasets needed for training reliable text classifiers. Recent research suggests using Large Language Models (LLMs) to automate the annotation…

Computation and Language · Computer Science 2025-01-27 Tomas Horych , Christoph Mandl , Terry Ruas , Andre Greiner-Petter , Bela Gipp , Akiko Aizawa , Timo Spinde

In this study, we investigate how language models develop preferences for \textit{idiomatic} as compared to \textit{linguistically acceptable} Swedish, both during pretraining and when adapting a model from English to Swedish. To do so, we…

Computation and Language · Computer Science 2026-02-04 Jenny Kunz

Detecting hate speech in videos remains challenging due to the complexity of multimodal content and the lack of fine-grained annotations in existing datasets. We present HateClipSeg, a large-scale multimodal dataset with both video-level…

Computer Vision and Pattern Recognition · Computer Science 2025-08-18 Han Wang , Zhuoran Wang , Roy Ka-Wei Lee

Large sense-annotated datasets are increasingly necessary for training deep supervised systems in Word Sense Disambiguation. However, gathering high-quality sense-annotated data for as many instances as possible is a laborious and expensive…

Computation and Language · Computer Science 2020-03-16 Tommaso Pasini , Jose Camacho-Collados

The surge in online content has created an urgent demand for robust detection systems, especially in non-English contexts where current tools demonstrate significant limitations. We present forePLay, a novel Polish language dataset for…

Computation and Language · Computer Science 2025-06-04 Anna Kołos , Katarzyna Lorenc , Emilia Wiśnios , Agnieszka Karlińska
‹ Prev 1 3 4 5 6 7 10 Next ›