English
Related papers

Related papers: FALL-E: A Foley Sound Synthesis Model and Strategi…

200 papers

Homomorphic encryption (HE) is a promising privacy-preserving technique for cross-silo federated learning (FL), where organizations perform collaborative model training on decentralized data. Despite the strong privacy guarantee, general HE…

Cryptography and Security · Computer Science 2021-09-30 Zhifeng Jiang , Wei Wang , Yang Liu

This paper describes the ESPnet-ST group's IWSLT 2021 submission in the offline speech translation track. This year we made various efforts on training data, architecture, and audio segmentation. On the data side, we investigated…

Audio and Speech Processing · Electrical Eng. & Systems 2021-07-07 Hirofumi Inaguma , Brian Yan , Siddharth Dalmia , Pengcheng Guo , Jiatong Shi , Kevin Duh , Shinji Watanabe

In this work, we propose a new recurrent autoencoder architecture, termed Feedback Recurrent AutoEncoder (FRAE), for online compression of sequential data with temporal dependency. The recurrent structure of FRAE is designed to efficiently…

Machine Learning · Computer Science 2020-02-18 Yang Yang , Guillaume Sautière , J. Jon Ryu , Taco S Cohen

Target speaker extraction (TSE) aims to isolate individual speaker voices from complex speech environments. The effectiveness of TSE systems is often compromised when the speaker characteristics are similar to each other. Recent research…

Sound · Computer Science 2024-10-08 Yun Liu , Xuechen Liu , Junichi Yamagishi

Recent advances in generative audio models have enabled high-fidelity environmental sound synthesis, raising serious concerns for audio security. The ESDD 2026 Challenge therefore addresses environmental sound deepfake detection under…

Sound · Computer Science 2025-12-24 Xiaoxuan Guo , Hengyan Huang , Jiayi Zhou , Renhe Sun , Jian Liu , Haonan Cheng , Long Ye , Qin Zhang

Many recent works on deep speaker embeddings train their feature extraction networks on large classification tasks, distinguishing between all speakers in a training set. Empirically, this has been shown to produce speaker-discriminative…

Sound · Computer Science 2020-02-04 Chau Luu , Peter Bell , Steve Renals

Deep generative models for audio synthesis have recently been significantly improved. However, the task of modeling raw-waveforms remains a difficult problem, especially for audio waveforms and music signals. Recently, the realtime audio…

Sound · Computer Science 2022-11-17 Seokjin Lee , Minhan Kim , Seunghyeon Shin , Daeho Lee , Inseon Jang , Wootaek Lim

We propose Serenade, a novel framework for the singing style conversion (SSC) task. Although singer identity conversion has made great strides in the previous years, converting the singing style of a singer has been an unexplored research…

Sound · Computer Science 2025-07-08 Lester Phillip Violeta , Wen-Chin Huang , Tomoki Toda

Sound design involves creatively selecting, recording, and editing sound effects for various media like cinema, video games, and virtual/augmented reality. One of the most time-consuming steps when designing sound is synchronizing audio…

This paper presents a circuit-algorithm co-design framework for learnable analog front-end (AFE) in audio signal classification. Designing AFE and backend classifiers separately is a common practice but non-ideal, as shown in this paper.…

Audio and Speech Processing · Electrical Eng. & Systems 2025-07-02 Jinhai Hu , Zhongyi Zhang , Cong Sheng Leow , Wang Ling Goh , Yuan Gao

Decompilation transforms compiled code back into a high-level programming language for analysis when source code is unavailable. Previous work has primarily focused on enhancing decompilation performance by increasing the scale of model…

Software Engineering · Computer Science 2024-10-04 Yunlong Feng , Dechuan Teng , Yang Xu , Honglin Mu , Xiao Xu , Libo Qin , Qingfu Zhu , Wanxiang Che

End-to-end (E2E) speech-to-text translation (ST) often depends on pretraining its encoder and/or decoder using source transcripts via speech recognition or text translation tasks, without which translation performance drops substantially.…

Computation and Language · Computer Science 2022-06-10 Biao Zhang , Barry Haddow , Rico Sennrich

Methods for modeling and controlling prosody with acoustic features have been proposed for neural text-to-speech (TTS) models. Prosodic speech can be generated by conditioning acoustic features. However, synthesized speech with a large…

Audio and Speech Processing · Electrical Eng. & Systems 2021-06-30 Taejun Bak , Jae-Sung Bae , Hanbin Bae , Young-Ik Kim , Hoon-Young Cho

Quantum Federated Learning (QFL) is an emerging paradigm that combines quantum computing and federated learning (FL) to enable decentralized model training while maintaining data privacy over quantum networks. However, quantum noise remains…

Quantum Physics · Physics 2025-07-18 Ratun Rahman , Atit Pokharel , Dinh C. Nguyen

In this technical report, the systems we submitted for subtask 4 of the DCASE 2021 challenge, regarding sound event detection, are described in detail. These models are closely related to the baseline provided for this problem, as they are…

Audio and Speech Processing · Electrical Eng. & Systems 2022-10-20 Wim Boes , Hugo Van hamme

In real dialogue scenarios, as there are unknown input noises in the utterances, existing supervised slot filling models often perform poorly in practical applications. Even though there are some studies on noise-robust models, these works…

Computation and Language · Computer Science 2023-10-06 Jiachi Liu , Liwen Wang , Guanting Dong , Xiaoshuai Song , Zechen Wang , Zhengyang Wang , Shanglin Lei , Jinzheng Zhao , Keqing He , Bo Xiao , Weiran Xu

Diffusion probabilistic models have been shown to generate state-of-the-art results on several competitive image synthesis benchmarks but lack a low-dimensional, interpretable latent space, and are slow at generation. On the other hand,…

Machine Learning · Computer Science 2022-11-30 Kushagra Pandey , Avideep Mukherjee , Piyush Rai , Abhishek Kumar

Most sound event detection (SED) systems perform well on clean datasets but degrade significantly in noisy environments. Language-queried audio source separation (LASS) models show promise for robust SED by separating target events;…

Sound · Computer Science 2025-08-12 Yuanjian Chen , Yang Xiao , Han Yin , Yadong Guan , Xubo Liu

In this work, we build a simple but strong baseline for sounding video generation. Given base diffusion models for audio and video, we integrate them with additional modules into a single model and train it to make the model jointly…

Machine Learning · Computer Science 2025-04-10 Masato Ishii , Akio Hayakawa , Takashi Shibuya , Yuki Mitsufuji

Environmental audio tagging aims to predict only the presence or absence of certain acoustic events in the interested acoustic scene. In this paper we make contributions to audio tagging in two parts, respectively, acoustic modeling and…

‹ Prev 1 8 9 10 Next ›