中文
相关论文

相关论文: Auto-Landmark: Acoustic Landmark Dataset and Open-…

200 篇论文

Furui first demonstrated that the identity of both consonant and vowel can be perceived from the C-V transition; later, Stevens proposed that acoustic landmarks are the primary cues for speech perception, and that steady-state regions are…

计算与语言 · 计算机科学 2018-05-16 Di He , Boon Pang Lim , Xuesong Yang , Mark Hasegawa-Johnson , Deming Chen

This paper introduces SoundSculpt, a neural network designed to extract target sound fields from ambisonic recordings. SoundSculpt employs an ambisonic-in-ambisonic-out architecture and is conditioned on both spatial information (e.g.,…

音频与语音处理 · 电气工程与系统科学 2025-06-03 Tuochao Chen , D Shin , Hakan Erdogan , Sinan Hersek

In a noisy environment, a lossy speech signal can be automatically restored by a listener if he/she knows the language well. That is, with the built-in knowledge of a "language model", a listener may effectively suppress noise interference…

机器学习 · 计算机科学 2019-07-03 Chien-Feng Liao , Yu Tsao , Xugang Lu , Hisashi Kawai

A dataset of anechoic recordings of various sound sources encountered in domestic environments is presented. The dataset is intended to be a resource of non-stationary, environmental noise signals that, when convolved with acoustic impulse…

音频与语音处理 · 电气工程与系统科学 2022-08-08 Philipp Götz , Cagdas Tuna , Andreas Walther , Emanuël A. P. Habets

Many ecosystems can undergo important qualitative changes, including sudden transitions to alternative stable states, in response to perturbations or increments in conditions. Such 'tipping points' are often preceded by declines in aspects…

种群与进化 · 定量生物学 2025-09-04 Neel P. Le Penru , Thomas M. Bury , Sarab S. Sethi , Robert M. Ewers , Lorenzo Picinali

Multi-modal learning in the audio-language domain has seen significant advancements in recent years. However, audio-language learning faces challenges due to limited and lower-quality data compared to image-language tasks. Existing…

音频与语音处理 · 电气工程与系统科学 2024-06-10 David Xu

Voice assistants increasingly rely on Speech Language Models (SpeechLMs) to interpret spoken queries and execute complex tasks, yet existing benchmarks lack domain breadth, acoustic diversity, and compositional reasoning complexity to…

Many datasets have been designed to further the development of fake audio detection. However, fake utterances in previous datasets are mostly generated by altering timbre, prosody, linguistic content or channel noise of original audio.…

Replay attacks remain a critical vulnerability for automatic speaker verification systems, particularly in real-time voice assistant applications. In this work, we propose acoustic maps as a novel spatial feature representation for replay…

音频与语音处理 · 电气工程与系统科学 2026-05-21 Michael Neri , Tuomas Virtanen

Audio-driven talking head generation requires precise synchronization between facial animations and audio signals. This paper introduces ATL-Diff, a novel approach addressing synchronization limitations while reducing noise and…

计算机视觉与模式识别 · 计算机科学 2025-07-18 Hoang-Son Vo , Quang-Vinh Nguyen , Seungwon Kim , Hyung-Jeong Yang , Soonja Yeom , Soo-Hyung Kim

Active Speaker Detection (ASD) aims to identify who is speaking in complex visual scenes. While humans naturally rely on lip-audio synchronization, existing ASD models often misclassify non-speaking instances when lip movements and audio…

计算机视觉与模式识别 · 计算机科学 2025-11-27 Le Thien Phuc Nguyen , Zhuoran Yu , Yong Jae Lee

This paper introduces Timers and Such, a new open source dataset of spoken English commands for common voice control use cases involving numbers. We describe the gap in existing spoken language understanding datasets that Timers and Such…

计算与语言 · 计算机科学 2021-10-04 Loren Lugosch , Piyush Papreja , Mirco Ravanelli , Abdelwahab Heba , Titouan Parcollet

This paper tests the hypothesis that distinctive feature classifiers anchored at phonetic landmarks can be transferred cross-lingually without loss of accuracy. Three consonant voicing classifiers were developed: (1) manually selected…

计算与语言 · 计算机科学 2017-08-23 Xiang Kong , Xuesong Yang , Mark Hasegawa-Johnson , Jeung-Yoon Choi , Stefanie Shattuck-Hufnagel

A judicious combination of dictionary learning methods, block sparsity and source recovery algorithm are used in a hierarchical manner to identify the noises and the speakers from a noisy conversation between two people. Conversations are…

声音 · 计算机科学 2016-10-31 K V Vijay Girish , A G Ramakrishnan , T V Ananthapadmanabha

Although speech recognition algorithms have developed quickly in recent years, achieving high transcription accuracy across diverse audio formats and acoustic environments remains a major challenge. This work explores how incorporating…

声音 · 计算机科学 2025-03-31 Aniket Abhishek Soni

Semantically-aligned $(speech, image)$ datasets can be used to explore "visually-grounded speech". In a majority of existing investigations, features of an image signal are extracted using neural networks "pre-trained" on other tasks (e.g.,…

机器学习 · 计算机科学 2020-10-30 Masood S. Mortazavi

In the speaker extraction problem, it is found that additional information from the target speaker contributes to the tracking and extraction of the target speaker, which includes voiceprint, lip movement, facial expression, and spatial…

音频与语音处理 · 电气工程与系统科学 2021-06-15 Yunzhe Hao , Jiaming Xu , Peng Zhang , Bo Xu

pyAMPACT (Python-based Automatic Music Performance Analysis and Comparison Toolkit) links symbolic and audio music representations to facilitate score-informed estimation of performance data in audio as well as general linking of symbolic…

声音 · 计算机科学 2026-01-06 Johanna Devaney , Daniel McKemie , Alex Morgan

Auditory attention to natural speech is a complex brain process. Its quantification from physiological signals can be valuable to improving and widening the range of applications of current brain-computer-interface systems, however it…

人机交互 · 计算机科学 2020-05-26 Nikesh Bajaj , Jesús Requena Carrión , Francesco Bellotti

With recent advancements in language technologies, humans are now speaking to devices. Increasing the reach of spoken language technologies requires building systems in local languages. A major bottleneck here are the underlying…

计算与语言 · 计算机科学 2021-02-23 Akshat Gupta , Xinjian Li , Sai Krishna Rallabandi , Alan W Black