中文
相关论文

相关论文: Binaural recording methods with analysis on inter-…

200 篇论文

Audio-driven facial reenactment is a crucial technique that has a range of applications in film-making, virtual avatars and video conferences. Existing works either employ explicit intermediate face representations (e.g., 2D facial…

计算机视觉与模式识别 · 计算机科学 2023-06-14 Ricong Huang , Peiwen Lai , Yipeng Qin , Guanbin Li

Lipreading has a lot of potential applications such as in the domain of surveillance and video conferencing. Despite this, most of the work in building lipreading systems has been limited to classifying silent videos into classes…

音频与语音处理 · 电气工程与系统科学 2019-07-03 Yaman Kumar , Rohit Jain , Khwaja Mohd. Salik , Rajiv Ratn Shah , Yifang yin , Roger Zimmermann

Connecting large libraries of digitized audio recordings to their corresponding sheet music images has long been a motivation for researchers to develop new cross-modal retrieval systems. In recent years, retrieval systems based on…

信息检索 · 计算机科学 2019-06-27 Stefan Balke , Matthias Dorfer , Luis Carvalho , Andreas Arzt , Gerhard Widmer

The dichotic method of hearing sound adapts in the region of musical harmony. The algorithm of the separation of the being dissonant voices into several separate groups is proposed. For an increase in the pleasantness of chords the…

声音 · 计算机科学 2024-05-09 Vadim R. Madgazin

We propose a neural network model that can separate target speech sources from interfering sources at different angular regions using two microphones. The model is trained with simulated room impulse responses (RIRs) using omni-directional…

音频与语音处理 · 电气工程与系统科学 2024-01-18 Yang Yang , George Sung , Shao-Fu Shih , Hakan Erdogan , Chehung Lee , Matthias Grundmann

Despite progress in video-to-audio generation, the field focuses predominantly on mono output, lacking spatial immersion. Existing binaural approaches remain constrained by a two-stage pipeline that first generates mono audio and then…

计算机视觉与模式识别 · 计算机科学 2025-12-03 Mengchen Zhang , Qi Chen , Tong Wu , Zihan Liu , Dahua Lin

This study examines pitch contours as a unifying semantic construct prevalent across various audio domains including music, speech, bioacoustics, and everyday sounds. Analyzing pitch contours offers insights into the universal role of pitch…

音频与语音处理 · 电气工程与系统科学 2025-03-26 Jakob Abeßer , Simon Schwär , Meinard Müller

Intraoperative ultrasound imaging provides real-time guidance during numerous surgical procedures, but its interpretation is complicated by noise, artifacts, and poor alignment with high-resolution preoperative MRI/CT scans. To bridge the…

计算机视觉与模式识别 · 计算机科学 2025-08-12 Noe Bertramo , Gabriel Duguey , Vivek Gopalakrishnan

Wearable technologies are envisaged to provide critical support to future healthcare systems. Hearables - devices worn in the ear - are of particular interest due to their ability to provide health monitoring in an efficient, reliable and…

Sound reconstruction via arbitrary objects has been a popular method in recent years, based on the recording of scattered light from the target object with a high-speed detector. In this work, we demonstrate the use of multi-mode fiber as a…

仪器与探测器 · 物理学 2024-05-06 Ege Küçükkömürcü , Berk Nezir Gün , Emre Yüce

Earable acoustic sensing offers a powerful and non-invasive modality for capturing fine-grained auditory and physiological signals directly from the ear canal, enabling continuous and context-aware monitoring of cognitive states. As earable…

人机交互 · 计算机科学 2025-12-23 Xijia Wei , Ting Dang , Khaldoon Al-Naimi , Yang Liu , Fahim Kawsar , Alessandro Montanari

Geometric perturbation theory is universal. A typical example is provided by the 3D wave equation, widely used in acoustics. We face vibrating eardrums as a binaural auditory input stemming from an external sound source. In the setup of…

数学物理 · 物理学 2020-03-18 David T. Heider , J. Leo van Hemmen

A number of auditory models have been developed using diverging approaches, either physiological or perceptual, but they share comparable stages of signal processing, as they are inspired by the same constitutive parts of the auditory…

音频与语音处理 · 电气工程与系统科学 2022-05-05 Alejandro Osses Vecchi , Léo Varnet , Laurel H. Carney , Torsten Dau , Ian C. Bruce , Sarah Verhulst , Piotr Majdak

High-fidelity binaural audio synthesis is crucial for immersive listening, but existing methods require extensive computational resources, limiting their edge-device application. To address this, we propose the Lightweight Implicit Neural…

音频与语音处理 · 电气工程与系统科学 2026-01-26 Xikun Lu , Fang Liu , Weizhi Shi , Jinqiu Sang

Pseudo-haptics exploit carefully crafted visual or auditory cues to trick the brain into "feeling" forces that are never physically applied, offering a low-cost alternative to traditional haptic hardware. Here, we present a comparative…

人机交互 · 计算机科学 2025-10-13 Nishant Gautam , Somya Sharma , Peter Corcoran , Kaspar Althoefer

Dubbing is a post-production process of re-recording actors' dialogues, which is extensively used in filmmaking and video production. It is usually performed manually by professional voice actors who read lines with proper prosody, and in…

音频与语音处理 · 电气工程与系统科学 2022-03-16 Chenxu Hu , Qiao Tian , Tingle Li , Yuping Wang , Yuxuan Wang , Hang Zhao

Ear acoustic authentication is a new biometrics method and it utilizes the differences in acoustic characteristics of the ear canal between users. However, there have been few reports on the factors that cause differences in the acoustic…

信号处理 · 电气工程与系统科学 2022-09-02 Riki Kimura , Shunsuke Tanaka , Naoki Wakui , Naoki Kodama , Shohei Yano

Given an arbitrary audio clip, audio-driven 3D facial animation aims to generate lifelike lip motions and facial expressions for a 3D head. Existing methods typically rely on training their models using limited public 3D datasets that…

计算机视觉与模式识别 · 计算机科学 2023-06-21 Liying Lu , Tianke Zhang , Yunfei Liu , Xuangeng Chu , Yu Li

From whirling ceiling fans to ticking clocks, the sounds that we hear subtly vary as we move through a scene. We ask whether these ambient sounds convey information about 3D scene structure and, if so, whether they provide a useful learning…

声音 · 计算机科学 2021-11-11 Ziyang Chen , Xixi Hu , Andrew Owens

Binaural rendering of ambisonic signals is of broad interest to virtual reality and immersive media. Conventional methods often require manually measured Head-Related Transfer Functions (HRTFs). To address this issue, we collect a paired…

声音 · 计算机科学 2022-11-07 Yin Zhu , Qiuqiang Kong , Junjie Shi , Shilei Liu , Xuzhou Ye , Ju-chiang Wang , Junping Zhang
‹ 上一页 1 8 9 10 下一页 ›