English
Related papers

Related papers: TorchAudio-Squim: Reference-less Speech Quality an…

200 papers

Recent work in the domain of speech enhancement has explored the use of self-supervised speech representations to aid in the training of neural speech enhancement models. However, much of this work focuses on using the deepest or final…

Sound · Computer Science 2023-06-27 George Close , William Ravenscroft , Thomas Hain , Stefan Goetze

Macroscopic intelligibility models predict the expected human word-error-rate for a given speech-in-noise stimulus. In contrast, microscopic intelligibility models aim to make fine-grained predictions about listeners' perception, e.g.…

Audio and Speech Processing · Electrical Eng. & Systems 2024-09-12 Paul Best , Santiago Cuervo , Ricard Marxer

Non-intrusive speech quality assessment (SQA) systems suffer from limited training data and costly human annotations, hindering their generalization to real-time conferencing calls. In this work, we propose leveraging large language models…

Audio and Speech Processing · Electrical Eng. & Systems 2025-08-11 Fredrik Cumlin , Xinyu Liang , Anubhab Ghosh , Saikat Chatterjee

Grammar competency estimation is essential for assessing linguistic proficiency in both written and spoken language; however, the spoken modality presents additional challenges due to its spontaneous, unstructured, and disfluent nature.…

Computation and Language · Computer Science 2025-11-18 Sourya Dipta Das , Shubham Kumar , Kuldeep Yadav

We consider the generative modeling of speech over multiple minutes, a requirement for long-form multimedia generation and audio-native voice assistants. However, textless spoken language models struggle to generate plausible speech past…

Computation and Language · Computer Science 2025-07-11 Se Jin Park , Julian Salazar , Aren Jansen , Keisuke Kinoshita , Yong Man Ro , RJ Skerry-Ryan

Large Language Models (LLMs) have demonstrated impressive capabilities across various domains, prompting a surge in their practical applications. However, concerns have arisen regarding the trustworthiness of LLMs outputs, particularly in…

Computation and Language · Computer Science 2024-05-08 Danna Zheng , Danyang Liu , Mirella Lapata , Jeff Z. Pan

Speech quality assessment is a problem for every researcher working on models that produce or process speech. Human subjective ratings, the gold standard in speech quality assessment, are expensive and time-consuming to acquire in a…

Sound · Computer Science 2023-05-25 Lorenz Diener , Marju Purin , Sten Sootla , Ando Saabas , Robert Aichner , Ross Cutler

In recent years, we have observed a rapid advancement in speech language models (SpeechLLMs), catching up with humans' listening and reasoning abilities. SpeechLLMs have demonstrated impressive spoken dialog question-answering (SQA)…

Computation and Language · Computer Science 2024-10-03 Junkai Wu , Xulin Fan , Bo-Ru Lu , Xilin Jiang , Nima Mesgarani , Mark Hasegawa-Johnson , Mari Ostendorf

Building on the success of large language models (LLMs), recent advancements such as GPT-4o have enabled real-time speech interactions through LLM-based voice assistants, offering a significantly improved user experience compared to…

Computation and Language · Computer Science 2024-12-12 Yiming Chen , Xianghu Yue , Chen Zhang , Xiaoxue Gao , Robby T. Tan , Haizhou Li

Advances in large language models (LLMs) have enabled significant capabilities in audio processing, resulting in state-of-the-art models now known as Large Audio Language Models (LALMs). However, minimal work has been done to measure audio…

Sound · Computer Science 2026-03-11 Laya Iyer , Angelina Wang , Sanmi Koyejo

Target speaker extraction (TSE) aims to recover a target speaker's speech from a mixture using a reference utterance as a cue. Most TSE systems adopt conditional auto-encoder architectures with one-step inference. Inspired by test-time…

Sound · Computer Science 2026-03-12 Zhenghai You , Ying Shi , Lantian Li , Dong Wang

This study proposes a multi-task pseudo-label learning (MPL)-based non-intrusive speech quality assessment model called MTQ-Net. MPL consists of two stages: obtaining pseudo-label scores from a pretrained model and performing multi-task…

Audio and Speech Processing · Electrical Eng. & Systems 2024-03-14 Ryandhimas E. Zezario , Bo-Ren Brian Bai , Chiou-Shann Fuh , Hsin-Min Wang , Yu Tsao

The detection of shouted speech is crucial in audio surveillance and monitoring. Although it is desirable for a security system to be able to identify emergencies, existing corpora provide only a binary label (i.e., shouted or normal) for…

Sound · Computer Science 2024-10-22 Takahiro Fukumori , Taito Ishida , Yoichi Yamashita

With the rapid integration of advanced reasoning capabilities into spoken dialogue models, the field urgently demands benchmarks that transcend simple interactions to address real-world complexity. However, current evaluations predominantly…

Computation and Language · Computer Science 2026-02-16 Yangzhuo Li , Shengpeng Ji , Yifu Chen , Tianle Liang , Haorong Ying , Yule Wang , Junbo Li , Jun Fang , Zhou Zhao

With the advances in speech communication systems such as online conferencing applications, we can seamlessly work with people regardless of where they are. However, during online meetings, speech quality can be significantly affected by…

Traditionally, the quality of acoustic echo cancellers is evaluated using intrusive speech quality assessment measures such as ERLE \cite{g168} and PESQ \cite{p862}, or by carrying out subjective laboratory tests. Unfortunately, the former…

Audio and Speech Processing · Electrical Eng. & Systems 2022-01-28 Marju Purin , Sten Sootla , Mateja Sponza , Ando Saabas , Ross Cutler

Cognitive impairment (CI) is of growing public health concern, and early detection is vital for effective intervention. Speech has gained attention as a non-invasive and easily collectible biomarker for assessing cognitive decline.…

Sound · Computer Science 2025-06-24 Mostafa Shahin , Beena Ahmed , Julien Epps

It is essential to perform speech intelligibility (SI) experiments with human listeners in order to evaluate objective intelligibility measures for developing effective speech enhancement and noise reduction algorithms. Recently,…

While speech Large Language Models (LLMs) excel at conventional tasks like basic speech recognition, they lack fine-grained, multi-dimensional perception. This deficiency is evident in their struggle to disentangle complex features like…

Audio and Speech Processing · Electrical Eng. & Systems 2026-05-13 Guojian Li , Zhixian Zhao , Zhennan Lin , Jingbin Hu , Qirui Zhan , Yuang Cao , Pengyuan Xie , Chuan Xie , Jie Liu , Qiang Zhang , Zhonghua Fu , Lei Xie

Spoken Language Understanding (SLU) is a task that aims to extract semantic information from spoken utterances. Previous research has made progress in end-to-end SLU by using paired speech-text data, such as pre-trained Automatic Speech…

Computation and Language · Computer Science 2023-07-11 Guan-Wei Wu , Guan-Ting Lin , Shang-Wen Li , Hung-yi Lee
‹ Prev 1 4 5 6 7 8 10 Next ›