English
Related papers

Related papers: Dialogue Enhancement in Object-based Audio -- Eval…

200 papers

Dialogue Enhancement (DE) enables the rebalancing of dialogue and background sounds to fit personal preferences and needs in the context of broadcast audio. When individual audio stems are unavailable from production, Dialogue Separation…

Audio and Speech Processing · Electrical Eng. & Systems 2023-06-01 Luca Resti , Martin Strauss , Matteo Torcoli , Emanuël Habets , Bernd Edler

Dialogue enhancement (DE) plays a vital role in broadcasting, enabling the personalization of the relative level between foreground speech and background music and effects. DE has been shown to improve the quality of experience,…

Audio and Speech Processing · Electrical Eng. & Systems 2023-03-24 Matteo Torcoli , Thomas Robotham , Emanuël A. P. Habets

Difficulties in following speech due to loud background sounds are common in broadcasting. Object-based audio, e.g., MPEG-H Audio solves this problem by providing a user-adjustable speech level. While object-based audio is gaining momentum,…

Dialog Enhancement (DE) is a feature which allows a user to increase the level of dialog in TV or movie content relative to non-dialog sounds. When only the original mix is available, DE is "unguided," and requires source separation. In…

Audio and Speech Processing · Electrical Eng. & Systems 2023-02-23 Aaron Master , Lie Lu , Jonas Samuelsson , Heidi-Maria Lehtonen , Scott Norcross , Nathan Swedlow , Audrey Howard

In TV services, dialogue level personalization is key to meeting user preferences and needs. When dialogue and background sounds are not separately available from the production stage, Dialogue Separation (DS) can estimate them to enable…

Audio and Speech Processing · Electrical Eng. & Systems 2023-03-24 Matteo Torcoli , Emanuël A. P. Habets

Blind and low-vision (BLV) people use audio descriptions (ADs) to access videos. However, current ADs are unalterable by end users, thus are incapable of supporting BLV individuals' potentially diverse needs and preferences. This research…

Human-Computer Interaction · Computer Science 2024-08-22 Rosiana Natalie , Ruei-Che Chang , Smitha Sheshadri , Anhong Guo , Kotaro Hara

LLM-based voice assistants (VAs) increasingly support older adults aging in place, yet how an assistant's agreeableness shapes explanation perception remains underexplored. We conducted a study(N=70) examining how VA agreeableness…

Human-Computer Interaction · Computer Science 2026-03-11 Niharika Mathur , Hasibur Rahman , Smit Desai

Since the advent of Deep Learning (DL), Speech Enhancement (SE) models have performed well under a variety of noise conditions. However, such systems may still introduce sonic artefacts, sound unnatural, and restrict the ability for a user…

Introduction: Virtual audiovisual technology and its methodology has yet to be established for psychoacoustic research. This study examined the effects of different audiovisual conditions on preference when listening to multi-talker…

Human-Computer Interaction · Computer Science 2023-01-18 Gerard Llorach , Maartje M. E. Hendrikse , Giso Grimm , Volker Hohmann

Existing deep learning (DL) based speech enhancement approaches are generally optimised to minimise the distance between clean and enhanced speech features. These often result in improved speech quality however they suffer from a lack of…

Sound · Computer Science 2021-11-19 Tassadaq Hussain , Mandar Gogate , Kia Dashtipour , Amir Hussain

Speech enhancement (SE) methods mainly focus on recovering clean speech from noisy input. In real-world speech communication, however, noises often exist in not only speaker but also listener environments. Although SE methods can suppress…

Audio and Speech Processing · Electrical Eng. & Systems 2022-03-23 Haoyu Li , Yun Liu , Junichi Yamagishi

Recent speech-to-speech (S2S) models generate intelligible speech but still lack natural expressiveness, largely due to the absence of a reliable evaluation metric. Existing approaches, such as subjective MOS ratings, low-level acoustic…

Sound · Computer Science 2025-10-24 Zhiyu Lin , Jingwen Yang , Jiale Zhao , Meng Liu , Sunzhu Li , Benyou Wang

Language Model (LM)-based speech enhancement (SE) has recently emerged as a promising direction, but existing approaches predominantly rely on token-level likelihood objectives that weakly reflect human perception. This mismatch limits…

Audio and Speech Processing · Electrical Eng. & Systems 2026-01-21 Haoyang Li , Nana Hou , Yuchen Hu , Jixun Yao , Sabato Marco Siniscalchi , Xuyi Zhuang , Deheng Ye , Wei Yang , Eng Siong Chng

High-tech Augmentative and Alternative Communication (AAC) has been rapidly advancing in recent years due to the increased use of large language models (LLMs) like ChatGPT, but many of these techniques are integrated without the inclusion…

Human-Computer Interaction · Computer Science 2025-08-06 Lara J. Martin , Malathy Nagalakshmi

Large Language Model-based Voice Assistants (LLM-VAs) are increasingly deployed in assistive settings for older adults, yet little is known about how an agent's personality shapes user perceptions of its explanations. This paper presents a…

Human-Computer Interaction · Computer Science 2026-04-30 Niharika Mathur , Hasibur Rahman , Smit Desai

While voice user interfaces offer increased accessibility due to hands-free and eyes-free interactions, older adults often have challenges such as constructing structured requests and perceiving how such devices operate. Voice-first user…

Human-Computer Interaction · Computer Science 2023-07-18 Chen Chen , Ella T. Lifset , Yichen Han , Arkajyoti Roy , Michael Hogarth , Alison A. Moore , Emilia Farcas , Nadir Weibel

Current hearing aids normally provide amplification based on a general prescriptive fitting, and the benefits provided by the hearing aids vary among different listening environments despite the inclusion of noise suppression feature.…

Sound · Computer Science 2021-06-10 Zehai Tu , Ning Ma , Jon Barker

Smart home automation systems aim to improve the comfort and convenience of users in their living environment. However, adapting automation to user needs remains a challenge. Indeed, many systems still rely on hand-crafted routines for each…

Human-Computer Interaction · Computer Science 2024-07-18 Jordan Rey-Jouanchicot , André Bottaro , Eric Campo , Jean-Léon Bouraoui , Nadine Vigouroux , Frédéric Vella

Older adults are using voice-based technologies in a variety of different contexts and are uniquely positioned to benefit from smart speakers' handsfree, voice-based interface. In order to better understand the ways in which older adults…

Human-Computer Interaction · Computer Science 2021-11-03 Margot Hanley , Shiri Azenkot

While current state-of-the-art Automatic Speech Recognition (ASR) systems achieve high accuracy on typical speech, they suffer from significant performance degradation on disordered speech and other atypical speech patterns. Personalization…

Audio and Speech Processing · Electrical Eng. & Systems 2021-06-21 Katrin Tomanek , Françoise Beaufays , Julie Cattiau , Angad Chandorkar , Khe Chai Sim
‹ Prev 1 2 3 10 Next ›