English
Related papers

Related papers: A Diffeomorphic Flow-based Variational Framework f…

200 papers

We propose a deep beamforming framework for enhancing target speaker(s) in multi-speaker environments. A deep neural network (DNN) is trained to estimate beamforming weights directly from noisy multichannel inputs while satisfying linear…

Audio and Speech Processing · Electrical Eng. & Systems 2026-05-21 Ilai Zaidel , Ori Engel , Bar Engel , Sharon Gannot

Generative flow networks (GFNs) are a class of models for sequential sampling of composite objects, which approximate a target distribution that is defined in terms of an energy function or a reward. GFNs are typically trained using a flow…

Machine Learning · Statistics 2022-10-17 Heiko Zimmermann , Fredrik Lindsten , Jan-Willem van de Meent , Christian A. Naesseth

Japan's kairanban culture and idobata conversations have long functioned as traditional communication practices that foster nuanced dialogue among community members and contribute to the formation of social balance. Inspired by these…

Computation and Language · Computer Science 2025-06-30 Takato Ueno , Keito Inoshita

In this paper, we propose a novel voice conversion strategy to resolve the mismatch between the training and conversion scenarios when parallel speech corpus is unavailable for training. Based on auto-encoder and disentanglement frameworks,…

Audio and Speech Processing · Electrical Eng. & Systems 2020-11-05 Yoohwan Kwon , Soo-Whan Chung , Hee-Soo Heo , Hong-Goo Kang

We propose a novel neural waveform compression method to catalyze emerging speech semantic communications. By introducing nonlinear transform and variational modeling, we effectively capture the dependencies within speech frames and…

Sound · Computer Science 2022-12-14 Shengshi Yao , Zixuan Xiao , Sixian Wang , Jincheng Dai , Kai Niu , Ping Zhang

We propose a Text-to-Speech method to create an unseen expressive style using one utterance of expressive speech of around one second. Specifically, we enhance the disentanglement capabilities of a state-of-the-art sequence-to-sequence…

Machine Learning · Computer Science 2020-02-18 Vatsal Aggarwal , Marius Cotescu , Nishant Prateek , Jaime Lorenzo-Trueba , Roberto Barra-Chicote

This paper focuses on using voice conversion (VC) to improve the speech intelligibility of surgical patients who have had parts of their articulators removed. Due to the difficulty of data collection, VC without parallel data is highly…

Audio and Speech Processing · Electrical Eng. & Systems 2019-08-26 Li-Wei Chen , Hung-Yi Lee , Yu Tsao

We introduce VocAlign, a novel source-free domain adaptation framework specifically designed for VLMs in open-vocabulary semantic segmentation. Our method adopts a student-teacher paradigm enhanced with a vocabulary alignment strategy,…

Computer Vision and Pattern Recognition · Computer Science 2025-09-30 Silvio Mazzucco , Carl Persson , Mattia Segu , Pier Luigi Dovesi , Federico Tombari , Luc Van Gool , Matteo Poggi

Despite the great promise of Transformers in many sequence modeling tasks (e.g., machine translation), their deterministic nature hinders them from generalizing to high entropy tasks such as dialogue response generation. Previous work…

Computation and Language · Computer Science 2020-03-31 Zhaojiang Lin , Genta Indra Winata , Peng Xu , Zihan Liu , Pascale Fung

Aligning Large Language Models (LLMs) with the diverse spectrum of human values remains a central challenge: preference-based methods often fail to capture deeper motivational principles. Value-based approaches offer a more principled path,…

Artificial Intelligence · Computer Science 2026-02-04 Woojin Kim , Sieun Hyeon , Jusang Oh , Jaeyoung Do

In this work, we address the problem of finegrained traceback of emotional and manipulation characteristics from synthetically manipulated speech. We hypothesize that combining semantic-prosodic cues captured by Speech Foundation Models…

Audio and Speech Processing · Electrical Eng. & Systems 2025-11-17 Girish , Mohd Mujtaba Akhtar , Farhan Sheth , Muskaan Singh

We propose an adversarial learning approach for generating multi-turn dialogue responses. Our proposed framework, hredGAN, is based on conditional generative adversarial networks (GANs). The GAN's generator is a modified hierarchical…

Computation and Language · Computer Science 2019-06-27 Oluwatobi Olabiyi , Alan Salimov , Anish Khazane , Erik T. Mueller

We propose Mask CycleGAN, a novel architecture for unpaired image domain translation built based on CycleGAN, with an aim to address two issues: 1) unimodality in image translation and 2) lack of interpretability of latent variables. Our…

Machine Learning · Computer Science 2022-05-17 Minfa Wang

In this paper, we present a latent variable (LV) framework to identify all the speakers and their keywords given a multi-speaker mixture signal. We introduce two separate LVs to denote active speakers and the keywords uttered. The…

Sound · Computer Science 2015-05-01 Harshavardhan Sundar , Thippur V. Sreenivas

The Emotional Voice Conversion (EVC) aims to convert the discrete emotional state from the source emotion to the target for a given speech utterance while preserving linguistic content. In this paper, we propose regularizing emotion…

Audio and Speech Processing · Electrical Eng. & Systems 2024-12-31 Ashishkumar Gudmalwar , Ishan D. Biyani , Nirmesh Shah , Pankaj Wasnik , Rajiv Ratn Shah

Variational encoder-decoders (VEDs) have shown promising results in dialogue generation. However, the latent variable distributions are usually approximated by a much simpler model than the powerful RNN structure used for encoding and…

Computation and Language · Computer Science 2018-02-07 Xiaoyu Shen , Hui Su , Shuzi Niu , Vera Demberg

Sentiment analysis of user-generated reviews or comments on products and services in social networks can help enterprises to analyze the feedback from customers and take corresponding actions for improvement. To mitigate large-scale…

Computation and Language · Computer Science 2021-02-19 Sicheng Zhao , Yang Xiao , Jiang Guo , Xiangyu Yue , Jufeng Yang , Ravi Krishna , Pengfei Xu , Kurt Keutzer

Recent neural networks such as WaveNet and sampleRNN that learn directly from speech waveform samples have achieved very high-quality synthetic speech in terms of both naturalness and speaker similarity even in multi-speaker text-to-speech…

Audio and Speech Processing · Electrical Eng. & Systems 2018-08-01 Yi Zhao , Shinji Takaki , Hieu-Thi Luong , Junichi Yamagishi , Daisuke Saito , Nobuaki Minematsu

Generative Flow Networks (GFlowNets) are amortized inference models designed to sample from unnormalized distributions over composable objects, with applications in generative modeling for tasks in fields such as causal discovery, NLP, and…

Machine Learning · Computer Science 2026-04-13 Tiago da Silva , Eliezer de Souza da Silva , Diego Mesquita

Emotional voice conversion (EVC) involves modifying various acoustic characteristics, such as pitch and spectral envelope, to match a desired emotional state while preserving the speaker's identity. Existing EVC methods often rely on text…

Sound · Computer Science 2025-01-22 Hyung-Seok Oh , Sang-Hoon Lee , Deok-Hyeon Cho , Seong-Whan Lee