English
Related papers

Related papers: ESPnet-EZ: Python-only ESPnet for Easy Fine-tuning…

200 papers

Different from the general visual classification, some classification tasks are more challenging as they need the professional categories of the images. In the paper, we call them expert-level classification. Previous fine-grained vision…

Computer Vision and Pattern Recognition · Computer Science 2022-11-30 Junde Wu , Huihui Fang , Yehui Yang , Yu Zhang , Haoyi Xiong , Huazhu Fu , Yanwu Xu

This paper describes the ESPnet-ST group's IWSLT 2021 submission in the offline speech translation track. This year we made various efforts on training data, architecture, and audio segmentation. On the data side, we investigated…

Audio and Speech Processing · Electrical Eng. & Systems 2021-07-07 Hirofumi Inaguma , Brian Yan , Siddharth Dalmia , Pengcheng Guo , Jiatong Shi , Kevin Duh , Shinji Watanabe

We present a state-of-the-art speech recognition system developed using end-to-end deep learning. Our architecture is significantly simpler than traditional speech systems, which rely on laboriously engineered processing pipelines; these…

Foundation models have shown great promise in speech emotion recognition (SER) by leveraging their pre-trained representations to capture emotion patterns in speech signals. To further enhance SER performance across various languages and…

Computation and Language · Computer Science 2024-06-18 Shahin Amiriparian , Filip Packań , Maurice Gerczuk , Björn W. Schuller

We present eCat, a novel end-to-end multispeaker model capable of: a) generating long-context speech with expressive and contextually appropriate prosody, and b) performing fine-grained prosody transfer between any pair of seen speakers.…

Audio and Speech Processing · Electrical Eng. & Systems 2023-06-21 Ammar Abbas , Sri Karlapati , Bastian Schnell , Penny Karanasou , Marcel Granero Moya , Amith Nagaraj , Ayman Boustati , Nicole Peinelt , Alexis Moinet , Thomas Drugman

Many recent studies have focused on fine-tuning pre-trained models for speech emotion recognition (SER), resulting in promising performance compared to traditional methods that rely largely on low-level, knowledge-inspired acoustic…

Sound · Computer Science 2024-02-15 Tiantian Feng , Shrikanth Narayanan

Speech signals in real-world environments are frequently affected by various distortions such as additive noise, reverberation, and bandwidth limitation, which may appear individually or in combination. Traditional speech enhancement…

Audio and Speech Processing · Electrical Eng. & Systems 2026-02-05 Doyeop Kwak , Youngjoon Jang , Seongyu Kim , Joon Son Chung

The presence of multiple talkers in the surrounding environment poses a difficult challenge for real-time speech communication systems considering the constraints on network size and complexity. In this paper, we present Personalized…

Audio and Speech Processing · Electrical Eng. & Systems 2021-06-09 Ritwik Giri , Shrikant Venkataramani , Jean-Marc Valin , Umut Isik , Arvindh Krishnaswamy

Recent approaches for few-shot 3D point cloud semantic segmentation typically require a two-stage learning process, i.e., a pre-training stage followed by a few-shot training stage. While effective, these methods face overreliance on…

Computer Vision and Pattern Recognition · Computer Science 2025-11-18 Jiahui Wang , Haiyue Zhu , Haoren Guo , Abdullah Al Mamun , Cheng Xiang , Tong Heng Lee

Over the past few years, speech enhancement methods based on deep learning have greatly surpassed traditional methods based on spectral subtraction and spectral estimation. Many of these new techniques operate directly in the the short-time…

Audio and Speech Processing · Electrical Eng. & Systems 2020-08-28 Jean-Marc Valin , Umut Isik , Neerad Phansalkar , Ritwik Giri , Karim Helwani , Arvindh Krishnaswamy

This paper presents EffortNet, a novel deep learning framework for decoding individual listening effort from electroencephalography (EEG) during speech comprehension. Listening effort represents a significant challenge in speech-hearing…

Audio and Speech Processing · Electrical Eng. & Systems 2025-08-22 Ching-Chih Sung , Cheng-Hung Hsin , Yu-Anne Shiah , Bo-Jyun Lin , Yi-Xuan Lai , Chia-Ying Lee , Yu-Te Wang , Borchin Su , Yu Tsao

Closed-Set speaker identification aims to assign a speech utterance to one of a predefined set of enrolled speakers and requires robust modeling of speaker-specific characteristics across multiple temporal scales. While recent deep learning…

Sound · Computer Science 2026-05-11 Yassin Terraf , Youssef Iraqi

The rise of deep learning has marked significant progress in fields such as computer vision, natural language processing, and medical imaging, primarily through the adaptation of pre-trained models for specific tasks. Traditional…

Machine Learning · Computer Science 2024-04-25 Charith Chandra Sai Balne , Sreyoshi Bhaduri , Tamoghna Roy , Vinija Jain , Aman Chadha

Electroencephalography (EEG) research typically focuses on tasks with narrowly defined objectives, but recent studies are expanding into the use of unlabeled data within larger models, aiming for a broader range of applications. This…

Signal Processing · Electrical Eng. & Systems 2025-05-26 Anders Gjølbye , Lina Skerath , William Lehn-Schiøler , Nicolas Langer , Lars Kai Hansen

The development of spiking neural network simulation software is a critical component enabling the modeling of neural systems and the development of biologically inspired algorithms. Existing software frameworks support a wide range of…

Neural and Evolutionary Computing · Computer Science 2019-03-27 Hananel Hazan , Daniel J. Saunders , Hassaan Khan , Darpan T. Sanghavi , Hava T. Siegelmann , Robert Kozma

Speech enhancement is a task to improve the intelligibility and perceptual quality of degraded speech signal. Recently, neural networks based methods have been applied to speech enhancement. However, many neural network based methods…

Sound · Computer Science 2021-02-22 Qiuqiang Kong , Haohe Liu , Xingjian Du , Li Chen , Rui Xia , Yuxuan Wang

SpeechPy is an open source Python package that contains speech preprocessing techniques, speech features, and important post-processing operations. It provides most frequent used speech features including MFCCs and filterbank energies…

Sound · Computer Science 2018-07-25 Amirsina Torfi

We propose Easy End-to-End Diffusion-based Text to Speech, a simple and efficient end-to-end text-to-speech model based on diffusion. E3 TTS directly takes plain text as input and generates an audio waveform through an iterative refinement…

Sound · Computer Science 2023-11-03 Yuan Gao , Nobuyuki Morioka , Yu Zhang , Nanxin Chen

In recent years, instruction tuning has gained increasing attention and emerged as a crucial technique to enhance the capabilities of Large Language Models (LLMs). To construct high-quality instruction datasets, many instruction processing…

Computation and Language · Computer Science 2024-06-25 Yixin Ou , Ningyu Zhang , Honghao Gui , Ziwen Xu , Shuofei Qiao , Yida Xue , Runnan Fang , Kangwei Liu , Lei Li , Zhen Bi , Guozhou Zheng , Huajun Chen

Syntactic dependencies can be predicted with high accuracy, and are useful for both machine-learned and pattern-based information extraction tasks. However, their utility can be improved. These syntactic dependencies are designed to…

Computation and Language · Computer Science 2020-06-05 Aryeh Tiktinsky , Yoav Goldberg , Reut Tsarfaty