中文
相关论文

相关论文: Leveraging Pre-Trained Autoencoders for Interpreta…

200 篇论文

Although prototypical network (ProtoNet) has proved to be an effective method for few-shot sound event detection, two problems still exist. Firstly, the small-scaled support set is insufficient so that the class prototypes may not represent…

声音 · 计算机科学 2022-06-07 Dongchao Yang , Helin Wang , Yuexian Zou , Zhongjie Ye , Wenwu Wang

Modelling the underlying mechanisms of neurodegenerative diseases demands methods that capture heterogeneous and spatially varying dynamics from sparse, high-dimensional neuroimaging data. Integrating partial differential equation (PDE)…

图像与视频处理 · 电气工程与系统科学 2025-09-19 Sanduni Pinnawala , Annabelle Hartanto , Ivor J. A. Simpson , Peter A. Wijeratne

Deep learning models have achieved high performance in medical applications, however, their adoption in clinical practice is hindered due to their black-box nature. Self-explainable models, like prototype-based models, can be especially…

计算机视觉与模式识别 · 计算机科学 2024-07-22 Shreyasi Pathak , Jörg Schlötterer , Jeroen Veltman , Jeroen Geerdink , Maurice van Keulen , Christin Seifert

Unsupervised disentangled representation learning from the unlabelled audio data, and high fidelity audio generation have become two linchpins in the machine learning research fields. However, the representation learned from an unsupervised…

音频与语音处理 · 电气工程与系统科学 2020-10-20 Kazi Nazmul Haque , Rajib Rana , Björn W Schuller

Recently, the success of pre-training in text domain has been fully extended to vision, audio, and cross-modal scenarios. The proposed pre-training models of different modalities are showing a rising trend of homogeneity in their model…

Automatic melody generation has been a long-time aspiration for both AI researchers and musicians. However, learning to generate euphonious melodies has turned out to be highly challenging. This paper introduces 1) a new variant of…

人工智能 · 计算机科学 2018-11-02 Yu-An Wang , Yu-Kai Huang , Tzu-Chuan Lin , Shang-Yu Su , Yun-Nung Chen

Modern LLMs face inference efficiency challenges due to their scale. To address this, many compression methods have been proposed, such as pruning and quantization. However, the effect of compression on a model's interpretability remains…

机器学习 · 计算机科学 2025-07-23 Suchit Gupte , Vishnu Kabir Chhabra , Mohammad Mahdi Khalili

Pattern recognition from audio signals is an active research topic encompassing audio tagging, acoustic scene classification, music classification, and other areas. Spectrogram and mel-frequency cepstral coefficients (MFCC) are among the…

音频与语音处理 · 电气工程与系统科学 2022-11-18 Md. Istiaq Ansari , Taufiq Hasan

High-dimensional data, particularly in the form of high-order tensors, presents a major challenge in self-supervised learning. While MLP-based autoencoders (AE) are commonly employed, their dependence on flattening operations exacerbates…

机器学习 · 计算机科学 2025-08-12 Junjing Zheng , Chengliang Song , Weidong Jiang , Xinyu Zhang

Visually grounded speech systems learn from paired images and their spoken captions. Recently, there have been attempts to utilize the visually grounded models trained from images and their corresponding text captions, such as CLIP, to…

音频与语音处理 · 电气工程与系统科学 2023-09-12 Saurabhchand Bhati , Jesús Villalba , Laureano Moro-Velazquez , Thomas Thebaud , Najim Dehak

To better support retrieval applications such as web search and question answering, growing effort is made to develop retrieval-oriented language models. Most of the existing works focus on improving the semantic representation capability…

计算与语言 · 计算机科学 2022-11-17 Shitao Xiao , Zheng Liu

We construct a new kind of encoder, leveraging the expressive power of diffusion models. In a traditional variational autoencoder, the encoder and decoder jointly negotiate a latent representation of the input. This is made possible by the…

机器学习 · 计算机科学 2026-05-14 Akhil Premkumar , Sarah Lucioni

The goal of universal audio representation learning is to obtain foundational models that can be used for a variety of downstream tasks involving speech, music and environmental sounds. To approach this problem, methods inspired by works on…

声音 · 计算机科学 2024-05-22 Leonardo Pepino , Pablo Riera , Luciana Ferrer

Recently, deep learning becomes the main focus of machine learning research and has greatly impacted many important fields. However, deep learning is criticized for lack of interpretability. As a successful unsupervised model in deep…

机器学习 · 计算机科学 2021-01-06 Fenglei Fan , Mengzhou Li , Yueyang Teng , Ge Wang

Variational Autoencoders for multimodal data hold promise for many tasks in data analysis, such as representation learning, conditional generation, and imputation. Current architectures either share the encoder output, decoder input, or…

In this paper we present a new approach to solve semi-supervised classification tasks for biomedical applications, involving a supervised autoencoder network. We create a network architecture that encodes labels into the latent space of an…

机器学习 · 计算机科学 2022-08-24 Cyprien Gille , Frederic Guyard , Michel Barlaud

Modeling and controlling complex spatiotemporal dynamical systems driven by partial differential equations (PDEs) often necessitate dimensionality reduction techniques to construct lower-order models for computational efficiency. This paper…

系统与控制 · 电气工程与系统科学 2024-09-12 Priyabrata Saha , Saibal Mukhopadhyay

We introduce a state-of-the-art real-time, high-fidelity, audio codec leveraging neural networks. It consists in a streaming encoder-decoder architecture with quantized latent space trained in an end-to-end fashion. We simplify and speed-up…

音频与语音处理 · 电气工程与系统科学 2022-10-25 Alexandre Défossez , Jade Copet , Gabriel Synnaeve , Yossi Adi

Deep neural networks have achieved remarkable performance in various text-based tasks but often lack interpretability, making them less suitable for applications where transparency is critical. To address this, we propose ProtoLens, a novel…

计算与语言 · 计算机科学 2024-10-25 Bowen Wei , Ziwei Zhu

In deep learning research, many melody extraction models rely on redesigning neural network architectures to improve performance. In this paper, we propose an input feature modification and a training objective modification based on two…

声音 · 计算机科学 2023-08-08 Keren Shao , Ke Chen , Taylor Berg-Kirkpatrick , Shlomo Dubnov