English
Related papers

Related papers: THUEE system description for NIST 2020 SRE CTS cha…

200 papers

This paper describes the system developed by the NPU team for the 2020 personalized voice trigger challenge. Our submitted system consists of two independently trained subsystems: a small footprint keyword spotting (KWS) system and a…

Sound · Computer Science 2021-03-01 Jingyong Hou , Li Zhang , Yihui Fu , Qing Wang , Zhanheng Yang , Qijie Shao , Lei Xie

Compared to hybrid automatic speech recognition (ASR) systems that use a modular architecture in which each component can be independently adapted to a new domain, recent end-to-end (E2E) ASR system are harder to customize due to their…

Computation and Language · Computer Science 2022-03-01 Samuel Thomas , Brian Kingsbury , George Saon , Hong-Kwang J. Kuo

As a practical alternative of speech separation, target speaker extraction (TSE) aims to extract the speech from the desired speaker using additional speaker cue extracted from the speaker. Its main challenge lies in how to properly extract…

Sound · Computer Science 2023-01-18 Kai Liu , Xucheng Wan , Ziqing Du , Huan Zhou

Although neural machine translation (NMT) has achieved impressive progress recently, it is usually trained on the clean parallel data set and hence cannot work well when the input sentence is the production of the automatic speech…

Computation and Language · Computer Science 2018-11-05 Xiang Li , Haiyang Xue , Wei Chen , Yang Liu , Yang Feng , Qun Liu

Target speaker extraction (TSE) aims to recover a target speaker's speech from a mixture using a reference utterance as a cue. Most TSE systems adopt conditional auto-encoder architectures with one-step inference. Inspired by test-time…

Sound · Computer Science 2026-03-12 Zhenghai You , Ying Shi , Lantian Li , Dong Wang

In recent years, the evolution of end-to-end (E2E) automatic speech recognition (ASR) models has been remarkable, largely due to advances in deep learning architectures like transformer. On top of E2E systems, researchers have achieved…

Audio and Speech Processing · Electrical Eng. & Systems 2024-06-12 Shiyi Han , Zhihong Lei , Mingbin Xu , Xingyu Na , Zhen Huang

In this paper, we describe our approach for the Podcast Summarisation challenge in TREC 2020. Given a podcast episode with its transcription, the goal is to generate a summary that captures the most important information in the content. Our…

Computation and Language · Computer Science 2021-01-14 Potsawee Manakul , Mark Gales

Recently there have been efforts to introduce new benchmark tasks for spoken language understanding (SLU), like semantic parsing. In this paper, we describe our proposed spoken semantic parsing system for the quality track (Track 1) in…

Computation and Language · Computer Science 2023-05-09 Siddhant Arora , Hayato Futami , Shih-Lun Wu , Jessica Huynh , Yifan Peng , Yosuke Kashiwagi , Emiru Tsunoo , Brian Yan , Shinji Watanabe

This paper describes our submission to the Second Clarity Enhancement Challenge (CEC2), which consists of target speech enhancement for hearing-aid (HA) devices in noisy-reverberant environments with multiple interferers such as music and…

Audio and Speech Processing · Electrical Eng. & Systems 2023-02-17 Samuele Cornell , Zhong-Qiu Wang , Yoshiki Masuyama , Shinji Watanabe , Manuel Pariente , Nobutaka Ono

In this paper, we present the system submission for the VoxCeleb Speaker Recognition Challenge 2020 (VoxSRC-20) by the DKU-DukeECE team. For track 1, we explore various kinds of state-of-the-art front-end extractors with different pooling…

Audio and Speech Processing · Electrical Eng. & Systems 2020-10-27 Weiqing Wang , Danwei Cai , Xiaoyi Qin , Ming Li

This paper describes the NPU-MSXF system for the IWSLT 2023 speech-to-speech translation (S2ST) task which aims to translate from English speech of multi-source to Chinese speech. The system is built in a cascaded manner consisting of…

Sound · Computer Science 2023-07-11 Kun Song , Yi lei , Peikun Chen , Yiqing Cao , Kun Wei , Yongmao Zhang , Lei Xie , Ning Jiang , Guoqing Zhao

This paper describes the ESPnet-ST group's IWSLT 2021 submission in the offline speech translation track. This year we made various efforts on training data, architecture, and audio segmentation. On the data side, we investigated…

Audio and Speech Processing · Electrical Eng. & Systems 2021-07-07 Hirofumi Inaguma , Brian Yan , Siddharth Dalmia , Pengcheng Guo , Jiatong Shi , Kevin Duh , Shinji Watanabe

Recent studies have made some progress in refining end-to-end (E2E) speech recognition encoders by applying Connectionist Temporal Classification (CTC) loss to enhance named entity recognition within transcriptions. However, these methods…

Audio and Speech Processing · Electrical Eng. & Systems 2023-12-12 Karan Singla , Shahab Jalalvand , Yeon-Jun Kim , Antonio Moreno Daniel , Srinivas Bangalore , Andrej Ljolje , Ben Stern

This technical report describes the system participating to the Detection and Classification of Acoustic Scenes and Events (DCASE) 2020 Challenge, Task 6: automated audio captioning. Our submission focuses on solving two indeterminacy…

Audio and Speech Processing · Electrical Eng. & Systems 2020-07-02 Yuma Koizumi , Daiki Takeuchi , Yasunori Ohishi , Noboru Harada , Kunio Kashino

Voice Assistants such as Alexa, Siri, and Google Assistant typically use a two-stage Spoken Language Understanding pipeline; first, an Automatic Speech Recognition (ASR) component to process customer speech and generate text transcriptions,…

Computation and Language · Computer Science 2020-12-17 Subendhu Rongali , Beiye Liu , Liwei Cai , Konstantine Arkoudas , Chengwei Su , Wael Hamza

We propose an end-to-end deep model for speaker verification in the wild. Our model uses thin-ResNet for extracting speaker embeddings from utterances and a Siamese capsule network and dynamic routing as the Back-end to calculate a…

Audio and Speech Processing · Electrical Eng. & Systems 2020-09-29 Amirhossein Hajavi , Ali Etemad

Current research in synthesized speech detection primarily focuses on the generalization of detection systems to unknown spoofing methods of noise-free speech. However, the performance of anti-spoofing countermeasures (CM) system is often…

Sound · Computer Science 2024-07-30 Yikang Wang , Xingming Wang , Hiromitsu Nishizaki , Ming Li

We present UTDUSS, the UTokyo-SaruLab system submitted to Interspeech2024 Speech Processing Using Discrete Speech Unit Challenge. The challenge focuses on using discrete speech unit learned from large speech corpora for some tasks. We…

Sound · Computer Science 2024-03-21 Wataru Nakata , Kazuki Yamauchi , Dong Yang , Hiroaki Hyodo , Yuki Saito

In this paper, we report our submitted system for the ZeroSpeech 2020 challenge on Track 2019. The main theme in this challenge is to build a speech synthesizer without any textual information or phonetic labels. In order to tackle those…

Computation and Language · Computer Science 2020-05-26 Andros Tjandra , Sakriani Sakti , Satoshi Nakamura

This paper presents the Voice Timbre Attribute Detection (vTAD) systems developed by the Digital Signal Processing & Speech Technology Laboratory (DSP&STL) of the Department of Electronic Engineering (EE) at The Chinese University of Hong…

Audio and Speech Processing · Electrical Eng. & Systems 2026-02-16 Aemon Yat Fei Chiu , Jingyu Li , Yusheng Tian , Guangyan Zhang , Tan Lee
‹ Prev 1 3 4 5 6 7 10 Next ›