English
Related papers

Related papers: Dialogue Enhancement in Object-based Audio -- Eval…

200 papers

We introduce a state-of-the-art audio-visual on-screen sound separation system which is capable of learning to separate sounds and associate them with on-screen objects by looking at in-the-wild videos. We identify limitations of previous…

Sound · Computer Science 2021-10-15 Efthymios Tzinis , Scott Wisdom , Tal Remez , John R. Hershey

Auditory attention and selective phase-locking are central to human speech understanding in complex acoustic scenes and cocktail party settings, yet these capabilities in multilingual subjects remain poorly understood. While machine…

Audio and Speech Processing · Electrical Eng. & Systems 2026-03-11 Sai Samrat Kankanala , Ram Chandra , Sriram Ganapathy

Data augmentation (DA) is crucial to mitigate model training instability and over-fitting problems in low-resource open-domain dialogue generation. However, traditional DA methods often neglect semantic data diversity, restricting the…

Computation and Language · Computer Science 2024-04-02 Zhenhua Liu , Tong Zhu , Jianxiang Xiang , Wenliang Chen

Removing background noise from speech audio has been the subject of considerable effort, especially in recent years due to the rise of virtual communication and amateur recordings. Yet background noise is not the only unpleasant disturbance…

Sound · Computer Science 2022-09-19 Joan Serrà , Santiago Pascual , Jordi Pons , R. Oguz Araz , Davide Scaini

Large Audio-Language Models (LALMs) can take audio and text as the inputs and answer questions about the audio. While prior LALMs have shown strong performance on standard benchmarks, there has been alarming evidence that LALMs can…

Audio and Speech Processing · Electrical Eng. & Systems 2025-09-16 Tzu-wen Hsu , Ke-Han Lu , Cheng-Han Chiang , Hung-yi Lee

This study proposes augmenting dialog data with think-aloud utterances (TAUs) for modeling individual personalities in text chat by LLM. TAU is a verbalization of a speaker's thought before articulating the utterance. We expect "persona…

Computation and Language · Computer Science 2025-10-30 Seiya Ishikura , Hiroaki Yamada , Tatsuya Hiraoka , Hiroaki Yamada , Takenobu Tokunaga

Large language models (LLMs) have shown great promise for capturing contextual information in natural language processing tasks. We propose a novel approach to speaker diarization that incorporates the prowess of LLMs to exploit contextual…

Audio and Speech Processing · Electrical Eng. & Systems 2023-09-15 Tae Jin Park , Kunal Dhawan , Nithin Koluguri , Jagadeesh Balam

Follow-up conversations with virtual assistants (VAs) enable a user to seamlessly interact with a VA without the need to repeatedly invoke it using a keyword (after the first query). Therefore, accurate Device-directed Speech Detection…

Audio and Speech Processing · Electrical Eng. & Systems 2024-11-06 Ognjen , Rudovic , Pranay Dighe , Yi Su , Vineet Garg , Sameer Dharur , Xiaochuan Niu , Ahmed H. Abdelaziz , Saurabh Adya , Ahmed Tewfik

The fast increase of web services and mobile apps, which collect personal data from users, increases the risk that their privacy may be severely compromised. In particular, the increasing variety of spoken language interfaces and voice…

Auditory attention decoding (AAD) is a technique used to identify and amplify the talker that a listener is focused on in a noisy environment. This is done by comparing the listener's brainwaves to a representation of all the sound sources…

Audio and Speech Processing · Electrical Eng. & Systems 2023-02-14 Cong Han , Vishal Choudhari , Yinghao Aaron Li , Nima Mesgarani

Searching for and making decisions about information is becoming increasingly difficult as the amount of information and number of choices increases. Recommendation systems help users find items of interest of a particular type, such as…

Information Retrieval · Computer Science 2011-07-04 M. H. Goker , P. Langley , C. A. Thompson

Recent advancements in Direct Preference Optimization (DPO) have significantly enhanced the alignment of Large Language Models (LLMs) with human preferences, owing to its simplicity and effectiveness. However, existing methods typically…

Computation and Language · Computer Science 2024-10-28 Shilong Li , Yancheng He , Hui Huang , Xingyuan Bu , Jiaheng Liu , Hangyu Guo , Weixun Wang , Jihao Gu , Wenbo Su , Bo Zheng

Augmented listening devices such as hearing aids often perform poorly in noisy and reverberant environments with many competing sound sources. Large distributed microphone arrays can improve performance, but data from remote microphones…

Audio and Speech Processing · Electrical Eng. & Systems 2019-12-12 Ryan M. Corey , Matthew D. Skarha , Andrew C. Singer

Spoken dialogue systems (SDSs) utilize automatic speech recognition (ASR) at the front end of their pipeline. The role of ASR in SDSs is to recognize information in user speech related to response generation appropriately. Examining…

Computation and Language · Computer Science 2025-10-09 Kiyotada Mori , Seiya Kawano , Chaoran Liu , Carlos Toshinori Ishi , Angel Fernando Garcia Contreras , Koichiro Yoshino

We investigate intelligent personal assistants (IPAs) accessibility for deaf and hard of hearing (DHH) people who can use their voice in everyday communication. The inability of IPAs to understand diverse accents including deaf speech…

Human-Computer Interaction · Computer Science 2026-01-23 Paige S. DeVries , Michaela Okosi , Ming Li , Nora Dunphy , Gidey Gezae , Dante Conway , Abraham Glasser , Raja Kushalnagar , Christian Vogler

Our objective is an audio-visual model for separating a single speaker from a mixture of sounds such as other speakers and background noise. Moreover, we wish to hear the speaker even when the visual cues are temporarily absent due to…

Computer Vision and Pattern Recognition · Computer Science 2019-07-12 Triantafyllos Afouras , Joon Son Chung , Andrew Zisserman

Omnimodal Large Language Models (OLLMs) have shown significant progress in integrating vision and text, but still struggle with integrating vision and audio, often exhibiting suboptimal performance when processing audio queries compared to…

Computer Vision and Pattern Recognition · Computer Science 2025-10-20 Rui Hu , Delai Qiu , Shuyu Wei , Jiaming Zhang , Yining Wang , Shengping Liu , Jitao Sang

Depression is one of the most prevalent mental health disorders globally. In recent years, multi-modal data, such as speech, video, and transcripts, has been increasingly used to develop AI-assisted depression assessment systems. Large…

A judicious combination of dictionary learning methods, block sparsity and source recovery algorithm are used in a hierarchical manner to identify the noises and the speakers from a noisy conversation between two people. Conversations are…

Sound · Computer Science 2016-10-31 K V Vijay Girish , A G Ramakrishnan , T V Ananthapadmanabha

Nowadays, speech is becoming a more common, if not standard, interface to technology. This can be seen in the trend of technology changes over the years. Increasingly, voice is used to control programs, appliances and personal devices…

Human-Computer Interaction · Computer Science 2019-09-10 Abraham Glasser