中文
相关论文

相关论文: Head Orientation Estimation with Distributed Micro…

200 篇论文

The ability to classify spoken speech based on the style of speaking is an important problem. With the advent of BPO's in recent times, specifically those that cater to a population other than the local population, it has become necessary…

计算与语言 · 计算机科学 2015-04-08 Sunil Kopparapu , Saurabh Bhatnagar , K. Sahana , Sathyanarayana , Akhilesh Srivastava , P. V. S. Rao

Estimating the head pose of a person is a crucial problem for numerous applications that is yet mainly addressed as a subtask of frontal pose prediction. We present a novel method for unconstrained end-to-end head pose estimation to tackle…

计算机视觉与模式识别 · 计算机科学 2023-09-15 Thorsten Hempel , Ahmed A. Abdelrahman , Ayoub Al-Hamadi

A Head Related Transfer Function (HRTF) characterizes how a human ear receives sounds from a point in space, and depends on the shapes of one's head, pinna, and torso. Accurate estimations of HRTFs for human subjects are crucial in enabling…

声音 · 计算机科学 2022-03-22 Navid H. Zandi , Awny M. El-Mohandes , Rong Zheng

Six-dimensional movable antenna (6DMA) is an innovative technology to improve wireless network capacity by adjusting 3D positions and 3D rotations of antenna surfaces based on channel spatial distribution. However, the existing works on…

信号处理 · 电气工程与系统科学 2024-12-09 Xiaodan Shao , Rui Zhang , Qijun Jiang , Jihong Park , Tony Q. S. Quek , Robert Schober

We consider the problem of audio voice separation for binaural applications, such as earphones and hearing aids. While today's neural networks perform remarkably well (separating $4+$ sources with 2 microphones) they assume a known or fixed…

声音 · 计算机科学 2022-07-18 Zhongweiyang Xu , Romit Roy Choudhury

We investigate a speech enhancement method based on the binaural coherence-to-diffuse power ratio (CDR), which preserves auditory spatial cues for maskers and a broadside target. Conventional CDR estimators typically rely on a mathematical…

音频与语音处理 · 电气工程与系统科学 2022-07-19 Reza Ghanavi , Craig Jin

This paper presents a sound source localization strategy that relies on a microphone array embedded in an unmanned ground vehicle and an asynchronous close-talking microphone near the operator. A signal coarse alignment strategy is combined…

机器人学 · 计算机科学 2025-07-30 Victor Liu , Timothy Du , Jordy Sehn , Jack Collier , François Grondin

Measuring personal head-related transfer functions (HRTFs) is essential in binaural audio. Personal HRTFs are not only required for binaural rendering and for loudspeaker-based binaural reproduction using crosstalk cancellation, but they…

音频与语音处理 · 电气工程与系统科学 2022-05-12 Tobias Kabzinski , Peter Jax

Human language is a combination of elemental languages/domains/styles that change across and sometimes within discourses. Language models, which play a crucial role in speech recognizers and machine translation systems, are particularly…

计算与语言 · 计算机科学 2013-03-22 Damianos Karakos , Mark Dredze , Sanjeev Khudanpur

This paper addresses the problem of binaural localization of a single speech source in noisy and reverberant environments. For a given binaural microphone setup, the binaural response corresponding to the direct-path propagation of a single…

声音 · 计算机科学 2016-09-08 Xiaofei Li , Laurent Girin , Radu Horaud , Sharon Gannot

The problem of source localization with ad hoc microphone networks in noisy and reverberant enclosures, given a training set of prerecorded measurements, is addressed in this paper. The training set is assumed to consist of a limited number…

声音 · 计算机科学 2016-10-18 Bracha Laufer-Goldshtein , Ronen Talmon , Sharon Gannot

Diagnosing autism spectrum disorder (ASD) by identifying abnormal speech patterns from examiner-patient dialogues presents significant challenges due to the subtle and diverse manifestations of speech-related symptoms in affected…

声音 · 计算机科学 2024-05-09 Chuanbo Hu , Jacob Thrasher , Wenqi Li , Mindi Ruan , Xiangxu Yu , Lynn K Paul , Shuo Wang , Xin Li

Auditory attention decoding (AAD) is a technique used to identify and amplify the talker that a listener is focused on in a noisy environment. This is done by comparing the listener's brainwaves to a representation of all the sound sources…

音频与语音处理 · 电气工程与系统科学 2023-02-14 Cong Han , Vishal Choudhari , Yinghao Aaron Li , Nima Mesgarani

We present a novel approach to the 3D sound source localization task for distributed ad-hoc microphone arrays by formulating it as a set-to-set regression problem. By training a multi-modal masked autoencoder model that operates on audio…

音频与语音处理 · 电气工程与系统科学 2024-12-17 Axel Berg , Jens Gulin , Mark O'Connor , Chuteng Zhou , Karl Åström , Magnus Oskarsson

Speech enhancement performance degrades significantly in noisy environments, limiting the deployment of speech-controlled technologies in industrial settings, such as manufacturing plants. Existing speech enhancement solutions primarly rely…

机器人学 · 计算机科学 2026-02-23 Zachary Turcotte , François Grondin

For robots to operate robustly in the real world, they should be aware of their uncertainty. However, most methods for object pose estimation return a single point estimate of the object's pose. In this work, we propose two learned methods…

计算机视觉与模式识别 · 计算机科学 2021-05-21 Brian Okorn , Mengyun Xu , Martial Hebert , David Held

Pose estimation refers to tracking a human's full body posture, including their head, torso, arms, and legs. The problem is challenging in practical settings where the number of body sensors are limited. Past work has shown promising…

计算机视觉与模式识别 · 计算机科学 2026-05-13 Sahil Bhandary Karnoor , Romit Roy Choudhury

The spatial information of sound plays a crucial role in various situations, ranging from daily activities to advanced engineering technologies. To fully utilize its potential, numerous research studies on spatial audio signal processing…

音频与语音处理 · 电气工程与系统科学 2025-03-14 Natsuki Ueno , Shoichi Koyama

Speaker identity plays a significant role in human communication and is being increasingly used in societal applications, many through advances in machine learning. Speaker identity perception is an essential cognitive phenomenon that can…

音频与语音处理 · 电气工程与系统科学 2024-06-18 Gasser Elbanna

The purpose of this work is to demonstrate a robust and clinically validated method for correcting sound speed aberrations in medical ultrasound. We propose a correction method that calculates focusing delays directly from the observed…