中文
相关论文

相关论文: UTMOS: UTokyo-SaruLab System for VoiceMOS Challeng…

200 篇论文

Contrastive, self-supervised learning (SSL) is used to train a model that predicts cancer type from miRNA, mRNA or RPPA expression data. This model, a pretrained FT-Transformer, is shown to outperform XGBoost and CatBoost, standard…

机器学习 · 计算机科学 2023-11-17 Christian John Hurry , Emma Slade

Estimating the perceived quality of an audio signal is critical for many multimedia and audio processing systems. Providers strive to offer optimal and reliable services in order to increase the user quality of experience (QoE). In this…

音频与语音处理 · 电气工程与系统科学 2019-03-19 Anderson R. Avila , Hannes Gamper , Chandan Reddy , Ross Cutler , Ivan Tashev , Johannes Gehrke

Multimodal emotion recognition is an important research topic in artificial intelligence. Over the past few decades, researchers have made remarkable progress by increasing the dataset size and building more effective algorithms. However,…

This paper introduces the system submitted by the Yidun NISP team to the video keyword wakeup challenge. We propose a mandarin keyword spotting system (KWS) with several novel and effective improvements, including a big backbone (B) model,…

计算与语言 · 计算机科学 2021-12-06 Yuting Yang , Binbin Du , Yingxin Zhang , Wenxuan Wang , Yuke Li

Background noise is a major source of quality impairments in Voice over Internet Protocol (VoIP) and Public Switched Telephone Network (PSTN) calls. Recent work shows the efficacy of deep learning for noise suppression, but the datasets…

This paper presents a speech intelligibility model based on automatic speech recognition (ASR), combining phoneme probabilities from deep neural networks (DNN) and a performance measure that estimates the word error rate from these…

Objective evaluation of synthesized speech is critical for advancing speech generation systems, yet existing metrics for intelligibility and prosody remain limited in scope and weakly correlated with human perception. Word Error Rate (WER)…

音频与语音处理 · 电气工程与系统科学 2026-01-21 Ismail Rasim Ulgen , Zongyang Du , Junchen Lu , Philipp Koehn , Berrak Sisman

This paper presents the multi-speaker multi-lingual few-shot voice cloning system developed by THU-HCSI team for LIMMITS'24 Challenge. To achieve high speaker similarity and naturalness in both mono-lingual and cross-lingual scenarios, we…

声音 · 计算机科学 2024-04-26 Yixuan Zhou , Shuoyi Zhou , Shun Lei , Zhiyong Wu , Menglin Wu

Self-supervised learning (SSL) approaches such as wav2vec 2.0 and HuBERT models have shown promising results in various downstream tasks in the speech community. In particular, speech representations learned by SSL models have been shown to…

音频与语音处理 · 电气工程与系统科学 2022-04-11 Eesung Kim , Jae-Jin Jeon , Hyeji Seo , Hoon Kim

Despite improvements in automatic speaker verification (ASV), vulnerability against spoofing attacks remains a major concern. In this study, we investigate the integration of ASV and countermeasure (CM) subsystems into a modular spoof-aware…

音频与语音处理 · 电气工程与系统科学 2025-09-17 Oguzhan Kurnaz , Tomi Kinnunen , Cemal Hanilci

In this paper, we describe the top-scoring submissions for team RTZR VoxCeleb Speaker Recognition Challenge 2022 (VoxSRC-22) in the closed dataset, speaker verification Track 1. The top performed system is a fusion of 7 models, which…

音频与语音处理 · 电气工程与系统科学 2022-09-22 Sangwon Suh , Sunjong Park

This paper presents the details of our system designed for the Task 1 of Multimodal Information Based Speech Processing (MISP) Challenge 2021. The purpose of Task 1 is to leverage both audio and video information to improve the…

This technical report describes the SJTU X-LANCE Lab system for the three tracks in CNSRC 2022. In this challenge, we explored the speaker embedding modeling ability of deep ResNet (Deeper r-vector). All the systems are only trained on the…

声音 · 计算机科学 2023-05-16 Zhengyang Chen , Bei Liu , Bing Han , Leying Zhang , Yanmin Qian

This paper introduces the model structure used in the SVDD 2024 Challenge. The SVDD 2024 challenge has been introduced this year for the first time. Singing voice deepfake detection (SVDD) which faces complexities due to informal speech…

声音 · 计算机科学 2024-10-03 Qishan Zhang , Shuangbing Wen , Fangke Yan , Tao Hu , Jun Li

Speech emotion recognition is a challenging classification task with natural emotional speech, especially when the distribution of emotion types is imbalanced in the training and test data. In this case, it is more difficult for a model to…

音频与语音处理 · 电气工程与系统科学 2024-05-31 Mingjie Chen , Hezhao Zhang , Yuanchao Li , Jiachen Luo , Wen Wu , Ziyang Ma , Peter Bell , Catherine Lai , Joshua Reiss , Lin Wang , Philip C. Woodland , Xie Chen , Huy Phan , Thomas Hain

This paper describes the BUCEA speaker diarization system for the 2022 VoxCeleb Speaker Recognition Challenge. Voxsrc-22 provides the development set and test set of VoxConverse, and we mainly use the test set of VoxConverse for parameter…

声音 · 计算机科学 2022-09-21 Ruohua Zhou , Yuxuan Du , Chenlei Hu

Many existing speaker verification systems are reported to be vulnerable against different spoofing attacks, for example speaker-adapted speech synthesis, voice conversion, play back, etc. In order to detect these spoofed speech signals as…

声音 · 计算机科学 2015-07-30 Shitao Weng , Shushan Chen , Lei Yu , Xuewei Wu , Weicheng Cai , Zhi Liu , Ming Li

In this technical report we describe the IDLAB top-scoring submissions for the VoxCeleb Speaker Recognition Challenge 2020 (VoxSRC-20) in the supervised and unsupervised speaker verification tracks. For the supervised verification tracks we…

音频与语音处理 · 电气工程与系统科学 2020-10-26 Jenthe Thienpondt , Brecht Desplanques , Kris Demuynck

Detecting out-of-distribution (OOD) inputs is a central challenge for safely deploying machine learning models in the real world. Existing solutions are mainly driven by small datasets, with low resolution and very few class labels (e.g.,…

计算机视觉与模式识别 · 计算机科学 2021-05-06 Rui Huang , Yixuan Li

This paper describes our proposed integration system for the spoofing-aware speaker verification challenge. It consists of a robust spoofing-aware verification system that use the speaker verification and antispoofing embeddings extracted…

音频与语音处理 · 电气工程与系统科学 2022-04-05 Juan M. Martín-Doñas , Iván G. Torre , Aitor Álvarez , Joaquin Arellano
‹ 上一页 1 8 9 10 下一页 ›