English
Related papers

Related papers: Enhancing Automatically Discovered Multi-level Aco…

200 papers

This paper is concerned with automatic continuous speech recognition using trainable systems. The aim of this work is to build acoustic models for spoken Swedish. This is done employing hidden Markov models and using the SpeechDat database…

Audio and Speech Processing · Electrical Eng. & Systems 2024-04-26 Giampiero Salvi

This research introduces a novel approach to textual and multimodal Hate Speech Detection (HSD), using Large Language Models (LLMs) as dynamic knowledge bases to generate background context and incorporate it into the input of HSD…

Computation and Language · Computer Science 2025-10-20 Joshua Wolfe Brook , Ilia Markov

Large language models record impressive performance on many natural language processing tasks. However, their knowledge capacity is limited to the pretraining corpus. Retrieval augmentation offers an effective solution by retrieving context…

Computation and Language · Computer Science 2023-11-22 Sai Munikoti , Anurag Acharya , Sridevi Wagle , Sameera Horawalavithana

Unsupervised spoken term discovery (UTD) aims at finding recurring segments of speech from a corpus of acoustic speech data. One potential approach to this problem is to use dynamic time warping (DTW) to find well-aligning patterns from the…

Audio and Speech Processing · Electrical Eng. & Systems 2020-08-04 Okko Räsänen , María Andrea Cruz Blandón

Multi-modal learning in the audio-language domain has seen significant advancements in recent years. However, audio-language learning faces challenges due to limited and lower-quality data compared to image-language tasks. Existing…

Audio and Speech Processing · Electrical Eng. & Systems 2024-06-10 David Xu

Transformative innovations in model architectures have introduced hierarchical embedding augmentation as a means to redefine the representation of tokens through multi-level semantic structures, offering enhanced adaptability to complex…

Computation and Language · Computer Science 2025-08-11 Derek Yotheringhay , Alistair Kirkland , Humphrey Kirkbride , Josiah Whitesteeple

Recent success of the Tacotron speech synthesis architecture and its variants in producing natural sounding multi-speaker synthesized speech has raised the exciting possibility of replacing expensive, manually transcribed, domain-specific,…

Computation and Language · Computer Science 2019-09-27 Andrew Rosenberg , Yu Zhang , Bhuvana Ramabhadran , Ye Jia , Pedro Moreno , Yonghui Wu , Zelin Wu

Automatic segmentation of text into minimal content-bearing units is an unsolved problem even for languages like English. Spaces between words offer an easy first approximation, but this approximation is not good enough for machine…

cmp-lg · Computer Science 2008-02-03 I. Dan Melamed

Topic classification systems on spoken documents usually consist of two modules: an automatic speech recognition (ASR) module to convert speech into text and a text topic classification (TTC) module to predict the topic class from the…

Computation and Language · Computer Science 2021-06-17 Tan Liu , Wu Guo , Bin Gu

Self-supervised learning has been used to leverage unlabelled data, improving accuracy and generalisation of speech systems through the training of representation models. While many recent works have sought to produce effective…

Computation and Language · Computer Science 2023-10-18 Antoni Dimitriadis , Siqi Pan , Vidhyasaharan Sethu , Beena Ahmed

Acoustic-to-Word recognition provides a straightforward solution to end-to-end speech recognition without needing external decoding, language model re-scoring or lexicon. While character-based models offer a natural solution to the…

Audio and Speech Processing · Electrical Eng. & Systems 2018-08-22 Shruti Palaskar , Florian Metze

In this paper, we propose a new strategy for acoustic scene classification (ASC) , namely recognizing acoustic scenes through identifying distinct sound events. This differs from existing strategies, which focus on characterizing global…

Sound · Computer Science 2019-10-23 Hongwei Song , Jiqing Han , Shiwen Deng , Zhihao Du

Expressive speech synthesis is crucial for many human-computer interaction scenarios, such as audiobooks, podcasts, and voice assistants. Previous works focus on predicting the style embeddings at one single scale from the information…

Sound · Computer Science 2023-08-01 Shun Lei , Yixuan Zhou , Liyang Chen , Zhiyong Wu , Xixin Wu , Shiyin Kang , Helen Meng

A mixed sample data augmentation strategy is proposed to enhance the performance of models on audio scene classification, sound event classification, and speech enhancement tasks. While there have been several augmentation methods shown to…

Sound · Computer Science 2021-08-09 Gwantae Kim , David K. Han , Hanseok Ko

Humans often speak in a continuous manner which leads to coherent and consistent prosody properties across neighboring utterances. However, most state-of-the-art speech synthesis systems only consider the information within each sentence…

Sound · Computer Science 2023-05-19 Ya-Jie Zhang , Wei Song , Yanghao Yue , Zhengchen Zhang , Youzheng Wu , Xiaodong He

Developing Text-to-Speech (TTS) systems that can synthesize natural breath is essential for human-like voice agents but requires extensive manual annotation of breath positions in training data. To this end, we propose a self-training…

Audio and Speech Processing · Electrical Eng. & Systems 2024-06-17 Dong Yang , Tomoki Koriyama , Yuki Saito

Automatic Pronunciation Assessment (APA) plays a vital role in Computer-assisted Pronunciation Training (CAPT) when evaluating a second language (L2) learner's speaking proficiency. However, an apparent downside of most de facto methods is…

Audio and Speech Processing · Electrical Eng. & Systems 2023-06-08 Fu-An Chao , Tien-Hong Lo , Tzu-I Wu , Yao-Ting Sung , Berlin Chen

Researches have shown accent classification can be improved by integrating semantic information into pure acoustic approach. In this work, we combine phonetic knowledge, such as vowels, with enhanced acoustic features to build an improved…

Sound · Computer Science 2016-02-25 Zhenhao Ge

We present Multiscale Audio Spectrogram Transformer (MAST) for audio classification, which brings the concept of multiscale feature hierarchies to the Audio Spectrogram Transformer (AST). Given an input audio spectrogram, we first patchify…

Audio and Speech Processing · Electrical Eng. & Systems 2023-05-19 Sreyan Ghosh , Ashish Seth , S. Umesh , Dinesh Manocha

Multimodal Stance Detection (MSD) is crucial for understanding public discourse, yet effectively fusing text and image, especially with conflicting signals, remains challenging. Existing methods often face difficulties with contextual…

Artificial Intelligence · Computer Science 2026-05-01 Weihai Lu , Zhejun Zhao , Yanshu Li , Huan He