中文
相关论文

相关论文: Head Orientation Estimation with Distributed Micro…

200 篇论文

In this paper we describe a speaker diarization system that enables localization and identification of all speakers present in a conversation or meeting. We propose a novel systematic approach to tackle several long-standing challenges in…

声音 · 计算机科学 2021-07-21 Siqi Zheng , Weilong Huang , Xianliang Wang , Hongbin Suo , Jinwei Feng , Zhijie Yan

Self-supervised speech representations are known to encode both speaker and phonetic information, but how they are distributed in the high-dimensional space remains largely unexplored. We hypothesize that they are encoded in orthogonal…

计算与语言 · 计算机科学 2023-12-12 Oli Liu , Hao Tang , Sharon Goldwater

Speech compression is commonly used to send voice over radio channels in applications such as mobile telephony and two-way push-to-talk (PTT) radio. In classical systems, the speech codec is combined with forward error correction,…

音频与语音处理 · 电气工程与系统科学 2025-07-29 David Rowe , Jean-Marc Valin

Speech signals, typically sampled at rates in the tens of thousands per second, contain redundancies, evoking inefficiencies in sequence modeling. High-dimensional speech features such as spectrograms are often used as the input for the…

We propose an audio-driven talking-head method to generate photo-realistic talking-head videos from a single reference image. In this work, we tackle two key challenges: (i) producing natural head motions that match speech prosody, and (ii)…

计算机视觉与模式识别 · 计算机科学 2021-07-21 Suzhen Wang , Lincheng Li , Yu Ding , Changjie Fan , Xin Yu

With millimeter wave wireless communications, the resulting radiation reflects on most visible objects, creating rich multipath environments, namely in urban scenarios. The radiation captured by a listening device is thus shaped by the…

信号处理 · 电气工程与系统科学 2018-04-12 João Gante , Gabriel Falcão , Leonel Sousa

Recent advancements in machine learning have significantly improved speech recognition, but recognizing speech from non-fluent or accented speakers remains a challenge. Previous efforts, relying on rule-based pronunciation patterns, have…

计算与语言 · 计算机科学 2025-06-04 Anna Seo Gyeong Choi , Jonghyeon Park , Myungwoo Oh

Consider a microphone array, such as those present in Amazon Echos, conference phones, or self-driving cars. One of the goals of these arrays is to decode the angles in which acoustic signals arrive at them. This paper considers the problem…

声音 · 计算机科学 2021-09-28 Yu-Lin Wei , Romit Roy Choudhury

Negative transfer in training of acoustic models for automatic speech recognition has been reported in several contexts such as domain change or speaker characteristics. This paper proposes a novel technique to overcome negative transfer by…

机器学习 · 计算机科学 2015-09-18 Mortaza Doulaty , Oscar Saz , Thomas Hain

Animating 3D head meshes using audio inputs has significant applications in AR/VR, gaming, and entertainment through 3D avatars. However, bridging the modality gap between speech signals and facial dynamics remains a challenge, often…

图形学 · 计算机科学 2025-05-13 Xinmu Wang , Xiang Gao , Xiyun Song , Heather Yu , Zongfang Lin , Liang Peng , Xianfeng Gu

In challenging real-life conditions such as extreme head-pose, occlusions, and low-resolution images where the visual information fails to estimate visual attention/gaze direction, audio signals could provide important and complementary…

计算机视觉与模式识别 · 计算机科学 2022-08-15 Shreya Ghosh , Abhinav Dhall , Munawar Hayat , Jarrod Knibbe

Reconstructing the room transfer functions needed to calculate the complex sound field in a room has several important real-world applications. However, an unpractical number of microphones is often required. Recently, in addition to…

音频与语音处理 · 电气工程与系统科学 2024-06-12 Francesca Ronchini , Luca Comanducci , Mirco Pezzoli , Fabio Antonacci , Augusto Sarti

Multi-channel speech separation using speaker's directional information has demonstrated significant gains over blind speech separation. However, it has two limitations. First, substantial performance degradation is observed when the coming…

声音 · 计算机科学 2023-02-28 Rongzhi Gu , Shi-Xiong Zhang , Dong Yu

This paper addresses the challenging scenario for the distant-talking control of a music playback device, a common portable speaker with four small loudspeakers in close proximity to one microphone. The user controls the device through…

声音 · 计算机科学 2014-05-07 Ramin Pichevar , Jason Wung , Daniele Giacobello , Joshua Atkins

Distributed transmit beamforming enables cooperative radios to act as one virtual antenna array, extending their communications' range beyond the capabilities of a single radio. Most existing distributed beamforming approaches rely on the…

信号处理 · 电气工程与系统科学 2022-01-14 Samer Hanna , Enes Krijestorac , Danijela Cabric

Conventional static measurement of head-related impulse responses (HRIRs) is time-consuming due to the need for repositioning a speaker array for each azimuth angle. Dynamic approaches using analytical models with a continuously rotating…

音频与语音处理 · 电气工程与系统科学 2025-04-22 Byeong-Yun Ko , Deokki Min , Hyeonuk Nam , Yong-Hwa Park

Own voice pickup for hearables in noisy environments benefits from using both an outer and an in-ear microphone outside and inside the occluded ear. Due to environmental noise recorded at both microphones, and amplification of the own voice…

音频与语音处理 · 电气工程与系统科学 2025-08-20 Mattes Ohlenbusch , Christian Rollwage , Simon Doclo

We study transfer learning in convolutional network architectures applied to the task of recognizing audio, such as environmental sound events and speech commands. Our key finding is that not only is it possible to transfer representations…

声音 · 计算机科学 2017-10-24 Brian McMahan , Delip Rao

This work aims to predict channels in wireless communication systems based on noisy observations, utilizing sequence-to-sequence models with attention (Seq2Seq-attn) and transformer models. Both models are adapted from natural language…

机器学习 · 统计学 2025-09-05 Valentina Rizzello , Benedikt Böck , Michael Joham , Wolfgang Utschick

This paper proposes PickNet, a neural network model for real-time channel selection for an ad hoc microphone array consisting of multiple recording devices like cell phones. Assuming at most one person to be vocally active at each time…

音频与语音处理 · 电气工程与系统科学 2022-01-25 Takuya Yoshioka , Xiaofei Wang , Dongmei Wang