English
Related papers

Related papers: MultiGraSCCo: A Multilingual Anonymization Benchma…

200 papers

High-quality preference data is essential for aligning foundation models with human values through preference learning. However, manual annotation of such data is often time-consuming and costly. Recent methods often adopt a self-rewarding…

Typical personal medical data contains sensitive information about individuals. Storing or sharing the personal medical data is thus often risky. For example, a short DNA sequence can provide information that can not only identify an…

Cryptography and Security · Computer Science 2019-02-01 Ho Bae , Dahuin Jung , Sungroh Yoon

In our previous work, we proposed a language-independent speaker anonymization system based on self-supervised learning models. Although the system can anonymize speech data of any language, the anonymization was imperfect, and the speech…

Sound · Computer Science 2022-03-29 Xiaoxiao Miao , Xin Wang , Erica Cooper , Junichi Yamagishi , Natalia Tomashenko

Machine learning (ML) algorithms are heavily based on the availability of training data, which, depending on the domain, often includes sensitive information about data providers. This raises critical privacy concerns. Anonymization…

Machine Learning · Computer Science 2025-11-03 Héber H. Arcolezi , Mina Alishahi , Adda-Akram Bendoukha , Nesrine Kaaniche

Stuttering is a complex disorder that requires specialized expertise for effective assessment and treatment. This paper presents an effort to enhance the FluencyBank dataset with a new stuttering annotation scheme based on established…

Ensuring native-like quality of large language model (LLM) responses across many languages is challenging. To address this, we introduce MENLO, a framework that operationalizes the evaluation of native-like response quality based on…

Computation and Language · Computer Science 2026-03-03 Chenxi Whitehouse , Sebastian Ruder , Tony Lin , Oksana Kurylo , Haruka Takagi , Janice Lam , Nicolò Busetto , Denise Diaz , Francisco Guzmán

Protecting privacy is essential when sharing data, particularly in the case of an online radicalization dataset that may contain personal information. In this paper, we explore the balance between preserving data usefulness and ensuring…

Computation and Language · Computer Science 2024-06-27 Arij Riabi , Menel Mahamdi , Virginie Mouilleron , Djamé Seddah

We compare three approaches to statistical machine translation (pure phrase-based, factored phrase-based and neural) by performing a fine-grained manual evaluation via error annotation of the systems' outputs. The error types in our…

Computation and Language · Computer Science 2018-02-13 Filip Klubička , Antonio Toral , Víctor M. Sánchez-Cartagena

Service providers of large language model (LLM) applications collect user instructions in the wild and use them in further aligning LLMs with users' intentions. These instructions, which potentially contain sensitive information, are…

Cryptography and Security · Computer Science 2024-07-03 Da Yu , Peter Kairouz , Sewoong Oh , Zheng Xu

Many under-resourced languages require high-quality datasets for specific tasks such as offensive language detection, disinformation, or misinformation identification. However, the intricacies of the content may have a detrimental effect on…

Computation and Language · Computer Science 2023-11-20 Stetsenko Daria

Speaker anonymization is the task of modifying a speech recording such that the original speaker cannot be identified anymore. Since the first Voice Privacy Challenge in 2020, along with the release of a framework, the popularity of this…

Sound · Computer Science 2023-12-25 Sarina Meyer , Xiaoxiao Miao , Ngoc Thang Vu

This work describes a self-supervised data augmentation approach used to improve learning models' performances when only a moderate amount of labeled data is available. Multiple copies of the original model are initially trained on the…

Computation and Language · Computer Science 2020-12-18 Gabriele Sarti

Synthetic data is often presented as a method for sharing sensitive information in a privacy-preserving manner by reproducing the global statistical properties of the original data without disclosing sensitive information about any…

Cryptography and Security · Computer Science 2022-11-22 Matteo Giomi , Franziska Boenisch , Christoph Wehmeyer , Borbála Tasnádi

The use of machine learning (ML)-based language models (LMs) to monitor content online is on the rise. For toxic text identification, task-specific fine-tuning of these models are performed using datasets labeled by annotators who provide…

Computation and Language · Computer Science 2021-12-08 Kofi Arhin , Ioana Baldini , Dennis Wei , Karthikeyan Natesan Ramamurthy , Moninder Singh

Incorporating every annotator's perspective is crucial for unbiased data modeling. Annotator fatigue and changing opinions over time can distort dataset annotations. To combat this, we propose to learn a more accurate representation of…

Machine Learning · Computer Science 2024-06-05 Uthman Jinadu , Yi Ding

This work proposes a novel privacy-preserving neural network feature representation to suppress the sensitive information of a learned space while maintaining the utility of the data. The new international regulation for personal data…

Computer Vision and Pattern Recognition · Computer Science 2020-08-10 Aythami Morales , Julian Fierrez , Ruben Vera-Rodriguez , Ruben Tolosana

Annotated data is an essential ingredient in natural language processing for training and evaluating machine learning models. It is therefore very desirable for the annotations to be of high quality. Recent work, however, has shown that…

Computation and Language · Computer Science 2022-09-27 Jan-Christoph Klie , Bonnie Webber , Iryna Gurevych

While the use of artificial intelligence (AI) for medical image analysis is gaining wide acceptance, the expertise, time and cost required to generate annotated data in the medical field are significantly high, due to limited availability…

Computer Vision and Pattern Recognition · Computer Science 2023-08-10 Abhishek Kushwaha , Sarthak Gupta , Anish Bhanushali , Tathagato Rai Dastidar

Datasets labelled by human annotators are widely used in the training and testing of machine learning models. In recent years, researchers are increasingly paying attention to label quality. However, it is not always possible to objectively…

Computer Vision and Pattern Recognition · Computer Science 2024-05-15 Luisa Schwirten , Jannes Scholz , Daniel Kondermann , Janis Keuper

As large language models (LLMs) rapidly advance and integrate into daily life, the privacy risks they pose are attracting increasing attention. We focus on a specific privacy risk where LLMs may help identify the authorship of anonymous…

Computation and Language · Computer Science 2024-11-21 Zichen Wen , Dadi Guo , Huishuai Zhang