English
Related papers

Related papers: ICAGC 2024: Inspirational and Convincing Audio Gen…

200 papers

The ISCSLP 2024 Conversational Voice Clone (CoVoC) Challenge aims to benchmark and advance zero-shot spontaneous style voice cloning, particularly focusing on generating spontaneous behaviors in conversational speech. The challenge…

This paper presents the NPU-HWC system submitted to the ISCSLP 2024 Inspirational and Convincing Audio Generation Challenge 2024 (ICAGC). Our system consists of two modules: a speech generator for Track 1 and a background audio generator…

Sound · Computer Science 2024-11-01 Dake Guo , Jixun Yao , Xinfa Zhu , Kangxiang Xia , Zhao Guo , Ziyu Zhang , Yao Wang , Jie Liu , Lei Xie

The ICASSP 2024 Speech Signal Improvement Grand Challenge is intended to stimulate research in the area of improving the speech signal quality in communication systems. This marks our second challenge, building upon the success from the…

The rapid advancement of Artificial Intelligence Generated Content (AIGC) technology has propelled audio-driven talking head generation, gaining considerable research attention for practical applications. However, performance evaluation…

Computer Vision and Pattern Recognition · Computer Science 2024-03-12 Weixia Zhang , Chengguang Zhu , Jingnan Gao , Yichao Yan , Guangtao Zhai , Xiaokang Yang

The ICASSP 2023 Acoustic Echo Cancellation Challenge is intended to stimulate research in acoustic echo cancellation (AEC), which is an important area of speech enhancement and is still a top issue in audio communication. This is the fourth…

Audio and Speech Processing · Electrical Eng. & Systems 2023-09-25 Ross Cutler , Ando Saabas , Tanel Parnamaa , Marju Purin , Evgenii Indenbom , Nicolae-Catalin Ristea , Jegor Gužvin , Hannes Gamper , Sebastian Braun , Robert Aichner

The Helsinki Speech Challenge 2024 (HSC2024) invites researchers to enhance and deconvolve speech audio recordings. We recorded a dataset that challenges participants to apply speech enhancement and inverse problems techniques to recorded…

Audio and Speech Processing · Electrical Eng. & Systems 2024-06-07 Martin Ludvigsen , Elli Karvonen , Markus Juvonen , Samuli Siltanen

This is the summary paper for the AudioMOS Challenge 2025, the very first challenge for automatic subjective quality prediction for synthetic audio. The challenge consists of three tracks. The first track aims to assess text-to-music…

Despite significant advancements in neural text-to-audio generation, challenges persist in controllability and evaluation. This paper addresses these issues through the Sound Scene Synthesis challenge held as part of the Detection and…

This paper reviews the AIS 2024 Video Quality Assessment (VQA) Challenge, focused on User-Generated Content (UGC). The aim of this challenge is to gather deep learning-based methods capable of estimating the perceptual quality of UGC…

The ICASSP 2023 Speech Signal Improvement Challenge is intended to stimulate research in the area of improving the speech signal quality in communication systems. The speech signal quality can be measured with SIG in ITU-T P.835 and is…

Audio and Speech Processing · Electrical Eng. & Systems 2023-10-16 Ross Cutler , Ando Saabas , Babak Naderi , Nicolae-Cătălin Ristea , Sebastian Braun , Solomiya Branets

Digital storytelling, as an art form, has struggled with cost-quality balance. The emergence of AI-generated Content (AIGC) is considered as a potential solution for efficient digital storytelling production. However, the specific form,…

Human-Computer Interaction · Computer Science 2023-09-29 Rongzhang Gu , Hui Li , Changyue Su , Wayne Wu

Real-world speech communication is rarely affected by a single type of degradation. Instead, it suffers from a complex interplay of acoustic interference, codec compression, and, increasingly, secondary artifacts introduced by upstream…

Sound · Computer Science 2025-12-30 Junan Zhang , Mengyao Zhu , Xin Xu , Hui Bu , Zhenhua Ling , Zhizheng Wu

The ICASSP 2022 Acoustic Echo Cancellation Challenge is intended to stimulate research in acoustic echo cancellation (AEC), which is an important area of speech enhancement and still a top issue in audio communication. This is the third AEC…

Audio and Speech Processing · Electrical Eng. & Systems 2022-03-01 Ross Cutler , Ando Saabas , Tanel Parnamaa , Marju Purin , Hannes Gamper , Sebastian Braun , Karsten Sørensen , Robert Aichner

Non-parallel many-to-many voice conversion remains an interesting but challenging speech processing task. Recently, AutoVC, a conditional autoencoder based method, achieved excellent conversion results by disentangling the speaker identity…

Sound · Computer Science 2022-08-09 Huaizhen Tang , Xulong Zhang , Jianzong Wang , Ning Cheng , Zhen Zeng , Edward Xiao , Jing Xiao

This paper describes the synthesis of the room acoustics challenge as a part of the generative data augmentation workshop at ICASSP 2025. The challenge defines a unique generative task that is designed to improve the quantity and diversity…

This paper reports on the NTIRE 2024 Quality Assessment of AI-Generated Content Challenge, which will be held in conjunction with the New Trends in Image Restoration and Enhancement Workshop (NTIRE) at CVPR 2024. This challenge is to…

Computer Vision and Pattern Recognition · Computer Science 2024-05-08 Xiaohong Liu , Xiongkuo Min , Guangtao Zhai , Chunyi Li , Tengchuan Kou , Wei Sun , Haoning Wu , Yixuan Gao , Yuqin Cao , Zicheng Zhang , Xiele Wu , Radu Timofte , Fei Peng , Huiyuan Fu , Anlong Ming , Chuanming Wang , Huadong Ma , Shuai He , Zifei Dou , Shu Chen , Huacong Zhang , Haiyi Xie , Chengwei Wang , Baoying Chen , Jishen Zeng , Jianquan Yang , Weigang Wang , Xi Fang , Xiaoxin Lv , Jun Yan , Tianwu Zhi , Yabin Zhang , Yaohui Li , Yang Li , Jingwen Xu , Jianzhao Liu , Yiting Liao , Junlin Li , Zihao Yu , Yiting Lu , Xin Li , Hossein Motamednia , S. Farhad Hosseini-Benvidi , Fengbin Guan , Ahmad Mahmoudi-Aznaveh , Azadeh Mansouri , Ganzorig Gankhuyag , Kihwan Yoon , Yifang Xu , Haotian Fan , Fangyuan Kong , Shiling Zhao , Weifeng Dong , Haibing Yin , Li Zhu , Zhiling Wang , Bingchen Huang , Avinab Saha , Sandeep Mishra , Shashank Gupta , Rajesh Sureddi , Oindrila Saha , Luigi Celona , Simone Bianco , Paolo Napoletano , Raimondo Schettini , Junfeng Yang , Jing Fu , Wei Zhang , Wenzhi Cao , Limei Liu , Han Peng , Weijun Yuan , Zhan Li , Yihang Cheng , Yifan Deng , Haohui Li , Bowen Qu , Yao Li , Shuqing Luo , Shunzhou Wang , Wei Gao , Zihao Lu , Marcos V. Conde , Xinrui Wang , Zhibo Chen , Ruling Liao , Yan Ye , Qiulin Wang , Bing Li , Zhaokun Zhou , Miao Geng , Rui Chen , Xin Tao , Xiaoyu Liang , Shangkun Sun , Xingyuan Ma , Jiaze Li , Mengduo Yang , Haoran Xu , Jie Zhou , Shiding Zhu , Bohan Yu , Pengfei Chen , Xinrui Xu , Jiabin Shen , Zhichao Duan , Erfan Asadi , Jiahe Liu , Qi Yan , Youran Qu , Xiaohui Zeng , Lele Wang , Renjie Liao

This paper presents Task 7 at the DCASE 2024 Challenge: sound scene synthesis. Recent advances in sound synthesis and generative models have enabled the creation of realistic and diverse audio content. We introduce a standardized evaluation…

Artificial Intelligence · Computer Science 2025-01-16 Mathieu Lagrange , Junwon Lee , Modan Tailleur , Laurie M. Heller , Keunwoo Choi , Brian McFee , Keisuke Imoto , Yuki Okamoto

This paper reports on the design and outcomes of the ICASSP SP Clarity Challenge: Speech Enhancement for Hearing Aids. The scenario was a listener attending to a target speaker in a noisy, domestic environment. There were multiple…

The ICASSP 2026 URGENT Challenge advances the series by focusing on universal speech enhancement (SE) systems that handle diverse distortions, domains, and input conditions. This overview paper details the challenge's motivation, task…

Audio and Speech Processing · Electrical Eng. & Systems 2026-01-21 Chenda Li , Wei Wang , Marvin Sach , Wangyou Zhang , Kohei Saijo , Samuele Cornell , Yihui Fu , Zhaoheng Ni , Tim Fingscheidt , Shinji Watanabe , Yanmin Qian

We present the third edition of the VoiceMOS Challenge, a scientific initiative designed to advance research into automatic prediction of human speech ratings. There were three tracks. The first track was on predicting the quality of…

‹ Prev 1 2 3 10 Next ›