English
Related papers

Related papers: DiffAnon: Diffusion-based Prosody Control for Voic…

200 papers

Recent advancements in diffusion models have revolutionized generative modeling. However, the impressive and vivid outputs they produce often come at the cost of significant model scaling and increased computational demands. Consequently,…

Machine Learning · Computer Science 2025-04-03 Jincheng Zhong , Xiangcheng Zhang , Jianmin Wang , Mingsheng Long

Privacy and security are major concerns when communicating speech signals to cloud services such as automatic speech recognition (ASR) and speech emotion recognition (SER). Existing solutions for speech anonymization mainly focus on voice…

Audio and Speech Processing · Electrical Eng. & Systems 2022-10-31 Minh Tran , Mohammad Soleymani

While large-scale pre-trained text-to-image models can synthesize diverse and high-quality human-centered images, novel challenges arise with a nuanced task of "identity fine editing": precisely modifying specific features of a subject…

Computer Vision and Pattern Recognition · Computer Science 2024-03-29 Haonan Lin , Mengmeng Wang , Yan Chen , Wenbin An , Yuzhe Yao , Guang Dai , Qianying Wang , Yong Liu , Jingdong Wang

Text-to-image diffusion models, such as Stable Diffusion, generate highly realistic images from text descriptions. However, the generation of certain content at such high quality raises concerns. A prominent issue is the accurate depiction…

Computer Vision and Pattern Recognition · Computer Science 2026-02-09 Liang Shi , Jie Zhang , Shiguang Shan

Toon shading is a type of non-photorealistic rendering task of animation. Its primary purpose is to render objects with a flat and stylized appearance. As diffusion models have ascended to the forefront of image synthesis methodologies,…

Computer Vision and Pattern Recognition · Computer Science 2024-01-30 Zhongjie Duan , Chengyu Wang , Cen Chen , Weining Qian , Jun Huang

Discrete audio tokens are compact representations that aim to preserve perceptual quality, phonetic content, and speaker characteristics while enabling efficient storage and inference, as well as competitive performance across diverse…

In an age of voice-enabled technology, voice anonymization offers a solution to protect people's privacy, provided these systems work equally well across subgroups. This study investigates bias in voice anonymization systems within the…

Audio and Speech Processing · Electrical Eng. & Systems 2023-11-28 Anna Leschanowsky , Ünal Ege Gaznepoglu , Nils Peters

Face stylization refers to the transformation of a face into a specific portrait style. However, current methods require the use of example-based adaptation approaches to fine-tune pre-trained generative models so that they demand lots of…

Computer Vision and Pattern Recognition · Computer Science 2023-12-06 Jin Liu , Huaibo Huang , Chao Jin , Ran He

We introduce DP-FinDiff, a differentially private diffusion framework for synthesizing mixed-type tabular data. DP-FinDiff employs embedding-based representations for categorical features, reducing encoding overhead and scaling to…

Machine Learning · Computer Science 2025-12-02 Timur Sattarov , Marco Schreyer , Damian Borth

Advances in talking-head animation based on Latent Diffusion Models (LDM) enable the creation of highly realistic, synchronized videos. These fabricated videos are indistinguishable from real ones, increasing the risk of potential misuse…

Graphics · Computer Science 2025-06-03 Yuan Gan , Jiaxu Miao , Yunze Wang , Yi Yang

Anonymous messaging platforms, such as Secret and Whisper, have emerged as important social media for sharing one's thoughts without the fear of being judged by friends, family, or the public. Further, such anonymous platforms are crucial…

Social and Information Networks · Computer Science 2015-04-28 Giulia Fanti , Peter Kairouz , Sewoong Oh , Pramod Viswanath

The increasingly pervasive facial recognition (FR) systems raise serious concerns about personal privacy, especially for billions of users who have publicly shared their photos on social media. Several attempts have been made to protect…

Computer Vision and Pattern Recognition · Computer Science 2025-12-02 Jiang Liu , Chun Pong Lau , Zhongliang Guo , Yuxiang Guo , Zhaoyang Wang , Rama Chellappa

In this paper, AN is introduced into semantic communication systems for the first time to prevent semantic eavesdropping. However, the introduction of AN also poses challenges for the legitimate receiver in extracting semantic information.…

Information Theory · Computer Science 2025-05-09 Boxiang He , Zihan Chen , Fanggang Wang , Shilian Wang , Zhijin Qin , Tony Q. S. Quek

Although voice conversion (VC) systems have shown a remarkable ability to transfer voice style, existing methods still have an inaccurate pitch and low speaker adaptation quality. To address these challenges, we introduce Diff-HierVC, a…

Audio and Speech Processing · Electrical Eng. & Systems 2023-11-09 Ha-Yeong Choi , Sang-Hoon Lee , Seong-Whan Lee

High-frequency gaze data contains more user-specific information than low-frequency data, promising for various applications. However, existing gaze modelling methods focus on low-frequency data, ignoring user-specific subtle eye movements…

Human-Computer Interaction · Computer Science 2025-05-21 Chuhan Jiao , Guanhua Zhang , Yeonjoo Cho , Zhiming Hu , Andreas Bulling

Voice anonymization techniques have been found to successfully obscure a speaker's acoustic identity in short, isolated utterances in benchmarks such as the VoicePrivacy Challenge. In practice, however, utterances seldom occur in isolation:…

Sound · Computer Science 2026-02-05 Cristina Aggazzotti , Ashi Garg , Zexin Cai , Nicholas Andrews

Audio inpainting seeks to restore missing segments in degraded recordings. Previous diffusion-based methods exhibit impaired performance when the missing region is large. We introduce the first approach that applies discrete diffusion over…

Sound · Computer Science 2026-02-18 Tali Dror , Iftach Shoham , Moshe Buchris , Oren Gal , Haim Permuter , Gilad Katz , Eliya Nachmani

We explore unsupervised speech enhancement using diffusion models as expressive generative priors for clean speech. Existing approaches guide the reverse diffusion process using noisy speech through an approximate, noise-perturbed…

Sound · Computer Science 2025-07-04 Mostafa Sadeghi , Jean-Eudes Ayilo , Romain Serizel , Xavier Alameda-Pineda

Although diffusion-based techniques have shown remarkable success in image generation and editing tasks, their abuse can lead to severe negative social impacts. Recently, some works have been proposed to provide defense against the abuse of…

Computer Vision and Pattern Recognition · Computer Science 2025-03-10 Zheng Li , Liangbin Xie , Jiantao Zhou , Xintao Wang , Haiwei Wu , Jinyu Tian

Prosody transfer is well-studied in the context of expressive speech synthesis. Cross-lingual prosody transfer, however, is challenging and has been under-explored to date. In this paper, we present a novel solution to learn prosody…

Audio and Speech Processing · Electrical Eng. & Systems 2023-06-21 Jakub Swiatkowski , Duo Wang , Mikolaj Babianski , Patrick Lumban Tobing , Ravichander Vipperla , Vincent Pollet
‹ Prev 1 4 5 6 7 8 10 Next ›