English
Related papers

Related papers: Non-locally averaged pruned reassigned spectrogram…

200 papers

In speech synthesis and speech enhancement systems, melspectrograms need to be precise in acoustic representations. However, the generated spectrograms are over-smooth, that could not produce high quality synthesized speech. Inspired by…

Audio and Speech Processing · Electrical Eng. & Systems 2019-12-04 Leyuan Sheng , Dong-Yan Huang , Evgeniy N. Pavlovskiy

Guided source separation (GSS) is a type of target-speaker extraction method that relies on pre-computed speaker activities and blind source separation to perform front-end enhancement of overlapped speech signals. It was first proposed…

Audio and Speech Processing · Electrical Eng. & Systems 2023-08-15 Desh Raj , Daniel Povey , Sanjeev Khudanpur

Explainability is a key component in many applications involving deep neural networks (DNNs). However, current explanation methods for DNNs commonly leave it to the human observer to distinguish relevant explanations from spurious noise.…

Machine Learning · Computer Science 2025-10-22 Paulo Yanez Sarmiento , Simon Witzke , Nadja Klein , Bernhard Y. Renard

Deep Learning based Automatic Speech Recognition (ASR) models are very successful, but hard to interpret. To gain better understanding of how Artificial Neural Networks (ANNs) accomplish their tasks, introspection methods have been…

Machine Learning · Computer Science 2020-02-20 Andreas Krug , Sebastian Stober

We introduce a novel method for Additive Noise Analysis for Persistence Thresholding (ANAPT) which separates significant features in the sublevel set persistence diagram of a time series based on a statistics analysis of the persistence of…

Algebraic Topology · Mathematics 2022-09-23 Audun D. Myers , Firas A. Khasawneh , Brittany T. Fasy

Panoptic Scene Graph Generation (PSG) parses objects and predicts their relationships (predicate) to connect human language and visual scenes. However, different language preferences of annotators and semantic overlaps between predicates…

Computer Vision and Pattern Recognition · Computer Science 2024-01-23 Li Li , Wei Ji , Yiming Wu , Mengze Li , You Qin , Lina Wei , Roger Zimmermann

Multi-resolution spectro-temporal features of a speech signal represent how the brain perceives sounds by tuning cortical cells to different spectral and temporal modulations. These features produce a higher dimensional representation of…

Audio and Speech Processing · Electrical Eng. & Systems 2022-06-28 Rahil Parikh , Nadee Seneviratne , Ganesh Sivaraman , Shihab Shamma , Carol Espy-Wilson

Filtered diode array spectrometers are routinely employed to infer the temporal evolution of spectral power from x-ray sources, but uniquely extracting spectral content from a finite set of broad, spectrally overlapping channel spectral…

Computational Physics · Physics 2020-08-03 G. E. Kemp , M. S. Rubery , C. D. Harris , M. J. May , K. Widmann , R. F. Heeter , S. B. Libby , M. B. Schneider , B. E. Blue

Frequency response function (FRF) estimation is a classical subject in system identification. In the past two decades, there have been remarkable advances in developing local methods for this subject, e.g., the local polynomial method,…

Systems and Control · Electrical Eng. & Systems 2024-12-31 Xiaozhu Fang , Yu Xu , Tianshi Chen

Large Language Models (LLMs) have shown improved generation performance through retrieval-augmented generation (RAG) following the retriever-reader paradigm, which supplements model inputs with externally retrieved knowledge. However, prior…

Computation and Language · Computer Science 2025-11-14 Zhanghao Hu , Qinglin Zhu , Siya Qi , Yulan He , Hanqi Yan , Lin Gui

Methods based on partial least squares (PLS) regression, which has recently gained much attention in the analysis of high-dimensional genomic datasets, have been developed since the early 2000s for performing variable selection. Most of…

Methodology · Statistics 2021-08-31 Jérémy Magnanensi , Myriam Maumy-Bertrand , Nicolas Meyer , Frédéric Bertrand

Foundation models and their checkpoints have significantly advanced deep learning, boosting performance across various applications. However, fine-tuned models often struggle outside their specific domains and exhibit considerable…

We introduce federated marginal personalization (FMP), a novel method for continuously updating personalized neural network language models (NNLMs) on private devices using federated learning (FL). Instead of fine-tuning the parameters of…

Computation and Language · Computer Science 2020-12-03 Zhe Liu , Fuchun Peng

Image signals typically are defined on a rectangular two-dimensional grid. However, there exist scenarios where this is not fulfilled and where the image information only is available for a non-regular subset of pixel position. For…

Image and Video Processing · Electrical Eng. & Systems 2022-07-15 Jürgen Seiler , André Kaup

This study evaluates the efficacy of two machine learning (ML) techniques, namely artificial neural networks (ANN) and gene expression programming (GEP) that use data-driven modeling to predict wall pressure spectra (WPS) underneath…

Fluid Dynamics · Physics 2024-02-27 Nachiketa Narayan Kurhade , Nagabhushana Rao Vadlamani , Akash Haridas

This paper proposes a unified deep speaker embedding framework for modeling speech data with different sampling rates. Considering the narrowband spectrogram as a sub-image of the wideband spectrogram, we tackle the joint modeling problem…

Audio and Speech Processing · Electrical Eng. & Systems 2020-12-02 Weicheng Cai , Ming Li

A typical audio signal processing pipeline includes multiple disjoint analysis stages, including calculation of a time-frequency representation followed by spectrogram-based feature analysis. We show how time-frequency analysis and…

Machine Learning · Statistics 2019-04-30 William J. Wilkinson , Michael Riis Andersen , Joshua D. Reiss , Dan Stowell , Arno Solin

We present a neural text-to-speech system for fine-grained prosody transfer from one speaker to another. Conventional approaches for end-to-end prosody transfer typically use either fixed-dimensional or variable-length prosody embedding via…

Audio and Speech Processing · Electrical Eng. & Systems 2019-07-05 Viacheslav Klimkov , Srikanth Ronanki , Jonas Rohnke , Thomas Drugman

With active research in audio compression techniques yielding substantial breakthroughs, spectral reconstruction of low-quality audio waves remains a less indulged topic. In this paper, we propose a novel approach for reconstructing higher…

Sound · Computer Science 2021-08-10 Darshan Deshpande , Harshavardhan Abichandani

With recent advances of diffusion model, generative speech enhancement (SE) has attracted a surge of research interest due to its great potential for unseen testing noises. However, existing efforts mainly focus on inherent properties of…

Audio and Speech Processing · Electrical Eng. & Systems 2024-06-05 Yuchen Hu , Chen Chen , Ruizhe Li , Qiushi Zhu , Eng Siong Chng
‹ Prev 1 4 5 6 7 8 10 Next ›