English
Related papers

Related papers: Enabling Automatic Self-Talk Detection via Earable…

200 papers

Human speech perception is multimodal. In natural speech, lip movements can precede corresponding voicing by a non-negligible gap of 100-300 ms, especially for specific consonants, affecting the time course of neural phonetic encoding in…

Audio and Speech Processing · Electrical Eng. & Systems 2025-06-26 Yi Wang , Oli Danyi Liu , Peter Bell

Sound morphing is the process of gradually and smoothly transforming one sound into another to generate novel and perceptually hybrid sounds that simultaneously resemble both. Recently, diffusion-based text-to-audio models have produced…

Audio and Speech Processing · Electrical Eng. & Systems 2024-08-15 Purnima Kamath , Chitralekha Gupta , Suranga Nanayakkara

Mental manipulation is a subtle yet pervasive form of abuse in interpersonal communication, making its detection critical for safeguarding potential victims. However, due to manipulation's nuanced and context-specific nature, identifying…

Stuttering is a common speech impediment that is caused by irregular disruptions in speech production, affecting over 70 million people across the world. Standard automatic speech processing tools do not take speech ailments into account…

Sound · Computer Science 2024-07-17 Liangyu Nie , Sudarsana Reddy Kadiri , Ruchit Agrawal

This paper addresses the problem of single-channel speech separation, where the number of speakers is unknown, and each speaker may speak multiple utterances. We propose a speech separation model that simultaneously performs separation,…

Audio and Speech Processing · Electrical Eng. & Systems 2025-05-23 Yuzhu Wang , Archontis Politis , Konstantinos Drossos , Tuomas Virtanen

Recognizing speaking in humans is a central task towards understanding social interactions. Ideally, speaking would be detected from individual voice recordings, as done previously for meeting scenarios. However, individual voice recordings…

Computer Vision and Pattern Recognition · Computer Science 2024-03-05 Jose Vargas Quiros , Chirag Raman , Stephanie Tan , Ekin Gedik , Laura Cabrera-Quiros , Hayley Hung

Neural TTS has shown it can generate high quality synthesized speech. In this paper, we investigate the multi-speaker latent space to improve neural TTS for adapting the system to new speakers with only several minutes of speech or…

Audio and Speech Processing · Electrical Eng. & Systems 2019-09-04 Yan Deng , Lei He , Frank Soong

End-to-end spoken dialogue models such as GPT-4o-audio have recently garnered significant attention in the speech domain. However, the evaluation of spoken dialogue models' conversational performance has largely been overlooked. This is…

Audio and Speech Processing · Electrical Eng. & Systems 2025-09-24 Shengpeng Ji , Tianle Liang , Yangzhuo Li , Jialong Zuo , Minghui Fang , Jinzheng He , Yifu Chen , Zhengqing Liu , Ziyue Jiang , Xize Cheng , Siqi Zheng , Jin Xu , Junyang Lin , Zhou Zhao

Lipreading has a lot of potential applications such as in the domain of surveillance and video conferencing. Despite this, most of the work in building lipreading systems has been limited to classifying silent videos into classes…

Audio and Speech Processing · Electrical Eng. & Systems 2019-07-03 Yaman Kumar , Rohit Jain , Khwaja Mohd. Salik , Rajiv Ratn Shah , Yifang yin , Roger Zimmermann

Human language can be expressed in either written or spoken form, i.e. text or speech. Humans can acquire knowledge from text to improve speaking and listening. However, the quest for speech pre-trained models to leverage unpaired text has…

Audio and Speech Processing · Electrical Eng. & Systems 2024-08-06 Duo Ma , Xianghu Yue , Junyi Ao , Xiaoxue Gao , Haizhou Li

This paper presents a macroscopic approach to automatic detection of speech sound disorder (SSD) in child speech. Typically, SSD is manifested by persistent articulation and phonological errors on specific phonemes in the language. The…

Audio and Speech Processing · Electrical Eng. & Systems 2022-06-30 Si-Ioi Ng , Cymie Wing-Yee Ng , Jiarui Wang , Tan Lee

We introduce a new automatic evaluation method for speaker similarity assessment, that is consistent with human perceptual scores. Modern neural text-to-speech models require a vast amount of clean training data, which is why many solutions…

Sound · Computer Science 2022-07-04 Deja Kamil , Sanchez Ariadna , Roth Julian , Cotescu Marius

Training large foundation models using self-supervised objectives on unlabeled data, followed by fine-tuning on downstream tasks, has emerged as a standard procedure. Unfortunately, the efficacy of this approach is often constrained by both…

We release MMSMR, a Massively Multi-System MultiReference dataset to enable future work on metrics and evaluation for dialog. Automatic metrics for dialogue evaluation should be robust proxies for human judgments; however, the verification…

Computation and Language · Computer Science 2024-11-20 Huda Khayrallah , Zuhaib Akhtar , Edward Cohen , Jyothir S , João Sedoc

Speaker diarization systems often struggle with high intrinsic intra-speaker variability, such as shifts in emotion, health, or content. This can cause segments from the same speaker to be misclassified as different individuals, for…

Audio and Speech Processing · Electrical Eng. & Systems 2025-09-19 Miseul Kim , Soo Jin Park , Kyungguen Byun , Hyeon-Kyeong Shin , Sunkuk Moon , Shuhua Zhang , Erik Visser

Large Language Models are increasingly used in conversational systems such as digital personal assistants, shaping how people interact with technology through language. While their responses often sound fluent and natural, they can also…

Computation and Language · Computer Science 2025-12-24 Heet Bodara , Md Masum Mushfiq , Isma Farah Siddiqui

In this study, we present SeMaScore, generated using a segment-wise mapping and scoring algorithm that serves as an evaluation metric for automatic speech recognition tasks. SeMaScore leverages both the error rate and a more robust…

Audio and Speech Processing · Electrical Eng. & Systems 2024-11-15 Zitha Sasindran , Harsha Yelchuri , T. V. Prabhakar

This paper addresses the challenging scenario for the distant-talking control of a music playback device, a common portable speaker with four small loudspeakers in close proximity to one microphone. The user controls the device through…

Sound · Computer Science 2014-05-07 Ramin Pichevar , Jason Wung , Daniele Giacobello , Joshua Atkins

Many promising applications of multimodal wearables require continuous sensing and heavy computation, yet users reject such devices due to privacy concerns. This paper shares our experiences building an ear-mounted voice-and-vision wearable…

Human-Computer Interaction · Computer Science 2025-11-26 Yonatan Tussa , Andy Heredia , Nirupam Roy

The rapid advancement of large language models (LLMs) has significantly propelled the development of text-based chatbots, demonstrating their capability to engage in coherent and contextually relevant dialogues. However, extending these…

Computation and Language · Computer Science 2024-08-23 Yinghao Aaron Li , Xilin Jiang , Jordan Darefsky , Ge Zhu , Nima Mesgarani
‹ Prev 1 3 4 5 6 7 10 Next ›