English
Related papers

Related papers: Statistical Speech Model Description with VMF Mixt…

200 papers

In this paper, we apply a latent class model (LCM) to the task of speaker diarization. LCM is similar to Patrick Kenny's variational Bayes (VB) method in that it uses soft information and avoids premature hard decisions in its iterations.…

Audio and Speech Processing · Electrical Eng. & Systems 2019-04-26 Liang He , Xianhong Chen , Can Xu , Yi Liu , Jia Liu , Michael T Johnson

Recent deep learning models have achieved high performance in speech enhancement; however, it is still challenging to obtain a fast and low-complexity model without significant performance degradation. Previous knowledge distillation…

Sound · Computer Science 2022-11-01 Wooseok Shin , Hyun Joon Park , Jin Sob Kim , Byung Hoon Lee , Sung Won Han

Diffusion model (DM) based Video Super-Resolution (VSR) approaches achieve impressive perceptual quality. However, they suffer from error accumulation, spatial artifacts, and a trade-off between perceptual quality and fidelity, primarily…

Computer Vision and Pattern Recognition · Computer Science 2026-03-30 Jingyi Xu , Meisong Zheng , Ying Chen , Minglang Qiao , Xin Deng , Mai Xu

Using data sources beyond the Automatic Identification System to represent the context a vessel is navigating in and consequently improve situation awareness is still rare in machine learning approaches to vessel trajectory prediction…

Machine Learning · Computer Science 2024-10-23 Kathrin Donandt , Dirk Söffker

Split federated learning (SFL) is a compute-efficient paradigm in distributed machine learning (ML), where components of large ML models are outsourced to remote servers. A significant challenge in SFL, particularly when deployed over…

Machine Learning · Computer Science 2025-10-28 Aladin Djuhera , Vlad C. Andrei , Xinyang Li , Ullrich J. Mönich , Holger Boche , Walid Saad

While Multimodal Large Language Models (MLLMs) have demonstrated impressive capabilities in general visual understanding, they frequently falter in fine-grained perception tasks that require identifying tiny objects or discerning subtle…

Computer Vision and Pattern Recognition · Computer Science 2026-04-15 Jilong Zhu , Yang Feng

We present a variational method for online state estimation and parameter learning in state-space models (SSMs), a ubiquitous class of latent variable models for sequential data. As per standard batch variational techniques, we use…

Machine Learning · Statistics 2022-06-16 Andrew Campbell , Yuyang Shi , Tom Rainforth , Arnaud Doucet

A promising approach for multi-microphone speech separation involves two deep neural networks (DNN), where the predicted target speech from the first DNN is used to compute signal statistics for time-invariant minimum variance…

Sound · Computer Science 2021-10-04 Zhong-Qiu Wang , Gordon Wichern , Jonathan Le Roux

We introduce Variational Latent Mode Decomposition (VLMD), a new algorithm for extracting oscillatory modes and associated connectivity structures from multivariate signals. VLMD addresses key limitations of existing Multivariate Mode…

Machine Learning · Computer Science 2025-05-26 Manuel Morante , Naveed ur Rehman

This paper proposes an approach to the joint modeling of the short-time Fourier transform magnitude and phase spectrograms with a deep generative model. We assume that the magnitude follows a Gaussian distribution and the phase follows a…

Sound · Computer Science 2022-07-18 Aditya Arie Nugraha , Kouhei Sekiguchi , Kazuyoshi Yoshii

This paper presents a novel unifying framework of bilinear LSTMs that can represent and utilize the nonlinear interaction of the input features present in sequence datasets for achieving superior performance over a linear LSTM and yet not…

Machine Learning · Computer Science 2023-09-12 Mohit Rajpal , Bryan Kian Hsiang Low

In the mobile communication field, some of the video applications boosted the interest of robust methods for video quality assessment. Out of all existing methods, We Preferred, No Reference Video Quality Assessment is the one which is most…

Computer Vision and Pattern Recognition · Computer Science 2023-03-10 Amitesh Kumar Singam , Benny Lövström , Wlodek J. Kulesza

A novel approach for speech segmentation is proposed, based on Multilevel Hybrid (mean/min) Filters (MHF) with the following features: An accurate transition location. Good performance in noisy environments (gaussian and impulsive noise).…

Audio and Speech Processing · Electrical Eng. & Systems 2022-03-04 Marcos Faundez-Zanuy , Francesc Vallverdu-Bayes

Vision Foundation Models (VFMs) and Vision-Language Models (VLMs) have gained traction in Domain Generalized Semantic Segmentation (DGSS) due to their strong generalization capabilities. However, existing DGSS methods often rely exclusively…

Computer Vision and Pattern Recognition · Computer Science 2025-04-16 Xin Zhang , Robby T. Tan

Probabilistic Latent Variable Models (LVMs) provide an alternative to self-supervised learning approaches for linguistic representation learning from speech. LVMs admit an intuitive probabilistic interpretation where the latent structure…

Audio and Speech Processing · Electrical Eng. & Systems 2020-09-09 Sameer Khurana , Antoine Laurent , Wei-Ning Hsu , Jan Chorowski , Adrian Lancucki , Ricard Marxer , James Glass

State-of-the-art language models (LMs) represented by long-short term memory recurrent neural networks (LSTM-RNNs) and Transformers are becoming increasingly complex and expensive for practical applications. Low-bit neural network…

Computation and Language · Computer Science 2021-12-22 Junhao Xu , Jianwei Yu , Shoukang Hu , Xunying Liu , Helen Meng

The Linear Parameter Varying Dynamical System (LPV-DS) is an effective approach that learns stable, time-invariant motion policies using statistical modeling and semi-definite optimization to encode complex motions for reactive robot…

Robotics · Computer Science 2024-03-26 Sunan Sun , Haihui Gao , Tianyu Li , Nadia Figueroa

With the development of deep learning, speech enhancement has been greatly optimized in terms of speech quality. Previous methods typically focus on the discriminative supervised learning or generative modeling, which tends to introduce…

Audio and Speech Processing · Electrical Eng. & Systems 2025-10-31 Nan Xu , Zhaolong Huang , Xiaonan Zhi

In this paper, we propose an adaptive framework for the variable power of the fractional least mean square (FLMS) algorithm. The proposed algorithm named as robust variable power FLMS (RVP-FLMS) dynamically adapts the fractional power of…

Optimization and Control · Mathematics 2017-02-07 Jawwad Ahmad , Muhammad Usman , Shujaat Khan , Imran Naseem , Hassan Jamil Syed

In this paper, we propose a new wireless video communication scheme to achieve high-efficiency video transmission over noisy channels. It exploits the idea of model division multiple access (MDMA) and extracts common semantic features…

Multimedia · Computer Science 2023-05-26 Zhicheng Bao , Haotai Liang , Chen Dong , Xiaodong Xu , Geng Liu