English
Related papers

Related papers: A comparative study of two-dimensional vocal tract…

200 papers

Self-supervised learning has demonstrated impressive performance in speech tasks, yet there remains ample opportunity for advancement in the realm of speech enhancement research. In addressing speech tasks, confining the attention mechanism…

Audio and Speech Processing · Electrical Eng. & Systems 2024-08-14 Tao Zheng , Liejun Wang , Yinfeng Yu

High-quality audio is essential in a wide range of applications, including online communication, virtual assistants, and the multimedia industry. However, degradation caused by noise, compression, and transmission artifacts remains a major…

This paper proposes a new channel estimation scheme for the multiuser massive multiple-input multiple-output (MIMO) systems in time-varying environment. We introduce a discrete Fourier transform (DFT) aided spatial-temporal basis expansion…

Information Theory · Computer Science 2016-11-01 Hongxiang Xie , Feifei Gao , Shun Zhang , Shi Jin

With the development of deep learning, speech enhancement has been greatly optimized in terms of speech quality. Previous methods typically focus on the discriminative supervised learning or generative modeling, which tends to introduce…

Audio and Speech Processing · Electrical Eng. & Systems 2025-10-31 Nan Xu , Zhaolong Huang , Xiaonan Zhi

Quality control in additive manufacturing (AM) is vital for industrial applications in areas such as the automotive, medical and aerospace sectors. Geometric inaccuracies caused by shrinkage and deformations can compromise the life and…

Computer Vision and Pattern Recognition · Computer Science 2025-10-24 Keerthana Chand , Tobias Fritsch , Bardia Hejazi , Konstantin Poka , Giovanni Bruno

This work aims to estimate time-resolved velocity field that is directly associated with pressure fluctuations in a subsonic round jet. To achieve this goal, synchronous measurements of the velocity field and in-flow pressure fluctuations…

Fluid Dynamics · Physics 2021-06-15 Songqi Li , Lawrence Ukeiley

Neural vocoder using denoising diffusion probabilistic model (DDPM) has been improved by adaptation of the diffusion noise distribution to given acoustic features. In this study, we propose SpecGrad that adapts the diffusion noise so that…

Audio and Speech Processing · Electrical Eng. & Systems 2022-08-08 Yuma Koizumi , Heiga Zen , Kohei Yatabe , Nanxin Chen , Michiel Bacchiani

Medical Image-to-image translation is a key task in computer vision and generative artificial intelligence, and it is highly applicable to medical image analysis. GAN-based methods are the mainstream image translation methods, but they…

Image and Video Processing · Electrical Eng. & Systems 2023-11-07 Zhuhui Wang , Jianwei Zuo , Xuliang Deng , Jiajia Luo

Vision Transformer (ViT) architectures represent images as collections of high-dimensional vectorized tokens, each corresponding to a rectangular non-overlapping patch. This representation trades spatial granularity for embedding…

Computer Vision and Pattern Recognition · Computer Science 2024-05-08 Dong Lao , Yangchao Wu , Tian Yu Liu , Alex Wong , Stefano Soatto

Providing quantitative interpretation of coherent nonlinear microscopy images, such as third-harmonic generation (THG), is generally hampered by the complex phase-matching conditions, especially in the presence of sample linear…

Optics · Physics 2026-03-09 Mohammad Reza Farhadinia , Nicolas Olivier

Investigating the sound field in and around ducts is an important topic in acoustics, e.g. when simulating musical instruments or the human vocal tract. In this paper a method that is based on the boundary element method in 3D combined with…

Numerical Analysis · Mathematics 2022-06-24 Wolfgang Kreuzer

An extremely fast time-harmonic finite element solver developed for the transmission analysis of photonic crystals was applied to mask simulation problems. The applicability was proven by examining a set of typical problems and by a…

Optics · Physics 2009-05-28 S. Burger , R. Köhle , L. Zschiedrich , W. Gao , F. Schmidt , R. März , C. Nölscher

The various speech sounds of a language are obtained by varying the shape and position of the articulators surrounding the vocal tract. Analyzing their variations is crucial for understanding speech production, diagnosing speech disorders…

Image and Video Processing · Electrical Eng. & Systems 2020-02-04 Mohammad Eslami , Christiane Neuschaefer-Rube , Antoine Serrurier

In recent years, speech diffusion models have advanced rapidly. Alongside the widely used U-Net architecture, transformer-based models such as the Diffusion Transformer (DiT) have also gained attention. However, current DiT speech models…

Understanding the relationship between vocal tract motion during speech and the resulting acoustic signal is crucial for aided clinical assessment and developing personalized treatment and rehabilitation strategies. Toward this goal, we…

Current deep neural network (DNN) based speech separation faces a fundamental challenge -- while the models need to be trained on short segments due to computational constraints, real-world applications typically require processing…

Audio and Speech Processing · Electrical Eng. & Systems 2025-07-04 Yuzhu Wang , Archontis Politis , Konstantinos Drossos , Tuomas Virtanen

Animal vocalizations contain sequential structures that carry important communicative information, yet most computational bioacoustics studies average the extracted frame-level features across the temporal axis, discarding the order of the…

Machine Learning · Computer Science 2025-11-14 Eklavya Sarkar , Mathew Magimai. -Doss

We describe here an experimental technique based on the acoustic scattering phenomenon allowing the direct probing of the vorticity field in a turbulent flow. Using time-frequency distributions, recently introduced in signal analysis…

chao-dyn · Physics 2009-10-31 Christophe Baudet , Olivier Michel , William J. Williams

Acoustic-to-articulatory inversion (AAI) methods estimate articulatory movements from the acoustic speech signal, which can be useful in several tasks such as speech recognition, synthesis, talking heads and language tutoring. Most earlier…

Audio and Speech Processing · Electrical Eng. & Systems 2020-08-06 Tamás Gábor Csapó

Sim2real transfer has received increasing attention lately due to the success of learning robotic tasks in simulation end-to-end. While there has been a lot of progress in transferring vision-based navigation policies, the existing sim2real…

Sound · Computer Science 2024-09-12 Changan Chen , Jordi Ramos , Anshul Tomar , Kristen Grauman
‹ Prev 1 8 9 10 Next ›