English
Related papers

Related papers: Non-locally averaged pruned reassigned spectrogram…

200 papers

Speech enhancement is a critical component of many user-oriented audio applications, yet current systems still suffer from distorted and unnatural outputs. While generative models have shown strong potential in speech synthesis, they are…

Audio and Speech Processing · Electrical Eng. & Systems 2022-02-11 Yen-Ju Lu , Zhong-Qiu Wang , Shinji Watanabe , Alexander Richard , Cheng Yu , Yu Tsao

We investigate the capability of a transformer pretrained on natural language to generalize to other modalities with minimal finetuning -- in particular, without finetuning of the self-attention and feedforward layers of the residual…

Machine Learning · Computer Science 2021-07-01 Kevin Lu , Aditya Grover , Pieter Abbeel , Igor Mordatch

Sparse Multi-view Images can be Learned to predict explicit radiance fields via Generalizable Gaussian Splatting approaches, which can achieve wider application prospects in real-life when ground-truth camera parameters are not required as…

Computer Vision and Pattern Recognition · Computer Science 2024-11-28 Yanyan Li , Yixin Fang , Federico Tombari , Gim Hee Lee

High resolution images can be acquired using a non-regular sampling sensor which consists of an underlying low resolution sensor that is covered with a non-regular sampling mask. The reconstructed high resolution image is then obtained…

Image and Video Processing · Electrical Eng. & Systems 2022-04-08 Markus Jonscher , Karina Jaskolka , Jürgen Seiler , André Kaup

In this paper, we investigate the application of graph signal processing (GSP) theory in speech enhancement. We first propose a set of shift operators to construct graph speech signals, and then analyze their spectrum in the graph Fourier…

Audio and Speech Processing · Electrical Eng. & Systems 2020-07-15 Xue Yan , Zhen Yang , Tingting Wang , Haiyan Guo

Three-dimensional Gaussian Splatting (3DGS) has recently emerged as an efficient representation for novel-view synthesis, achieving impressive visual quality. However, in scenes dominated by large and low-texture regions, common in indoor…

Computer Vision and Pattern Recognition · Computer Science 2025-10-29 Xirui Jin , Renbiao Jin , Boying Li , Danping Zou , Wenxian Yu

The possibility to tune the functional properties of nanomaterials is key to their technological applications. Superlattices, i.e., periodic repetitions of two or more materials in different dimensions are being explored for their potential…

Open-vocabulary panoptic reconstruction is a challenging task for simultaneous scene reconstruction and understanding. Recently, methods have been proposed for 3D scene understanding based on Gaussian splatting. However, these methods are…

Computer Vision and Pattern Recognition · Computer Science 2025-03-25 Yuxuan Xie , Xuan Yu , Changjian Jiang , Sitong Mao , Shunbo Zhou , Rui Fan , Rong Xiong , Yue Wang

Automatic speech recognition (ASR) for dysarthric speech remains challenging due to data scarcity, particularly in non-English languages. To address this, we fine-tune a voice conversion model on English dysarthric speech (UASpeech) to…

Linguistic information is encoded at varying timescales (subwords, phrases, etc.) and communicative levels, such as syntax and semantics. Contextualized embeddings have analogously been found to capture these phenomena at distinctive layers…

Computation and Language · Computer Science 2022-10-24 Max Müller-Eberstein , Rob van der Goot , Barbara Plank

Convolutional neural network (CNN) architectures have originated and revolutionized machine learning for images. In order to take advantage of CNNs in predictive modeling with audio data, standard FFT-based signal processing methods are…

Sound · Computer Science 2025-02-20 Pavol Harar , Roswitha Bammer , Anna Breger , Monika Dörfler , Zdenek Smekal

We investigate the ability to reconstruct and derive spatial structure from sparsely sampled 3D piezoresponse force microcopy data, captured using the band-excitation (BE) technique, via Gaussian Process (GP) methods. Even for weakly…

An innovative and novel quartz-enhanced photoacoustic spectroscopy (QEPAS) sensor for highly sensitive and selective breath gas analysis is introduced. The QEPAS sensor consists of two acoustically coupled micro-resonators (mR) with an…

Instrumentation and Detectors · Physics 2017-04-28 Jan C. Petersen , Laurent Lamard , Yuyang Feng , J-F. Focant , Andre peremans , Mikael Lassen

Historically, most speech models in machine-learning have used the mel-spectrogram as a speech representation. Recently, discrete audio tokens produced by neural audio codecs have become a popular alternate speech representation for speech…

Audio and Speech Processing · Electrical Eng. & Systems 2025-06-05 Ryan Langman , Ante Jukić , Kunal Dhawan , Nithin Rao Koluguri , Jason Li

Sparse-view tomographic reconstruction is a pivotal direction for reducing radiation dose and augmenting clinical applicability. While many research works have proposed the reconstruction of tomographic images from sparse 2D projections,…

Image and Video Processing · Electrical Eng. & Systems 2024-09-24 Jingmou Xian , Jian Zhu , Haolin Liao , Si Li

This paper proposes attentive statistics pooling for deep speaker embedding in text-independent speaker verification. In conventional speaker embedding, frame-level features are averaged over all the frames of a single utterance to form an…

Audio and Speech Processing · Electrical Eng. & Systems 2019-02-27 Koji Okabe , Takafumi Koshinaka , Koichi Shinoda

Recently, 3D Gaussian Splatting (3DGS) has become one of the mainstream methodologies for novel view synthesis (NVS) due to its high quality and fast rendering speed. However, as a point-based scene representation, 3DGS potentially…

Computer Vision and Pattern Recognition · Computer Science 2024-05-30 Zhaoliang Zhang , Tianchen Song , Yongjae Lee , Li Yang , Cheng Peng , Rama Chellappa , Deliang Fan

Cinematic audio source separation is a relatively new subtask of audio source separation, with the aim of extracting the dialogue, music, and effects stems from their mixture. In this work, we developed a model generalizing the Bandsplit…

Audio and Speech Processing · Electrical Eng. & Systems 2024-08-27 Karn N. Watcharasupat , Chih-Wei Wu , Yiwei Ding , Iroro Orife , Aaron J. Hipple , Phillip A. Williams , Scott Kramer , Alexander Lerch , William Wolcott

We introduce a new member of TSP (Time Stretched Pulse) for acoustic and speech measurement infrastructure, based on a simple all-pass filter and systematic randomization. This new infrastructure fundamentally upgrades our previous…

Sound · Computer Science 2021-02-16 Hideki Kawahara , Kohei Yatabe

A nanohertz-frequency stochastic gravitational-wave background can potentially be detected through the precise timing of an array of millisecond pulsars. This background produces low-frequency noise in the pulse arrival times that would…