English
Related papers

Related papers: Peransformer: Improving Low-informed Expressive Pe…

200 papers

This paper introduces an audio-visual speech enhancement system that leverages score-based generative models, also known as diffusion models, conditioned on visual information. In particular, we exploit audio-visual embeddings obtained from…

Audio and Speech Processing · Electrical Eng. & Systems 2023-06-05 Julius Richter , Simone Frintrop , Timo Gerkmann

In recent years, the burgeoning interest in diffusion models has led to significant advances in image and speech generation. Nevertheless, the direct synthesis of music waveforms from unrestricted textual prompts remains a relatively…

Sound · Computer Science 2023-09-22 Pengfei Zhu , Chao Pang , Yekun Chai , Lei Li , Shuohuan Wang , Yu Sun , Hao Tian , Hua Wu

This paper proposes an efficient memory transformer Emformer for low latency streaming speech recognition. In Emformer, the long-range history context is distilled into an augmented memory bank to reduce self-attention's computation…

Text-to-music (TTM) generation, which converts textual descriptions into audio, opens up innovative avenues for multimedia creation. Achieving high quality and diversity in this process demands extensive, high-quality data, which are often…

Sound · Computer Science 2025-06-18 Chang Li , Ruoyu Wang , Lijuan Liu , Jun Du , Yixuan Sun , Zilu Guo , Zhenrong Zhang , Yuan Jiang , Jianqing Gao , Feng Ma

Sortformer is an encoder-based speaker diarization model designed for supervising speaker tagging in speech-to-text models. Instead of relying solely on permutation invariant loss (PIL), Sortformer introduces Sort Loss to resolve the…

Audio and Speech Processing · Electrical Eng. & Systems 2025-07-22 Taejin Park , Ivan Medennikov , Kunal Dhawan , Weiqing Wang , He Huang , Nithin Rao Koluguri , Krishna C. Puvvada , Jagadeesh Balam , Boris Ginsburg

Recent illumination estimation methods have focused on enhancing the resolution and improving the quality and diversity of the generated textures. However, few have explored tailoring the neural network architecture to the Equirectangular…

Computer Vision and Pattern Recognition · Computer Science 2024-10-18 Jack Hilliard , Adrian Hilton , Jean-Yves Guillemaut

Recent advancements in speech encoders have drawn attention due to their integration with Large Language Models for various speech tasks. While most research has focused on either causal or full-context speech encoders, there's limited…

Lack of large, well-annotated emotional speech corpora continues to limit the performance and robustness of speech emotion recognition (SER), particularly as models grow more complex and the demand for multimodal systems increases. While…

Sound · Computer Science 2026-02-13 Chung-Soo Ahn , Rajib Rana , Sunil Sivadas , Carlos Busso , Jagath C. Rajapakse

Multi-level implicit discourse relation recognition (MIDRR) aims at identifying hierarchical discourse relations among arguments. Previous methods achieve the promotion through fine-tuning PLMs. However, due to the data scarcity and the…

Computation and Language · Computer Science 2024-02-26 Haodong Zhao , Ruifang He , Mengnan Xiao , Jing Xu

Transformer has obtained promising results on cognitive speech signal processing field, which is of interest in various applications ranging from emotion to neurocognitive disorder analysis. However, most works treat speech signal as a…

Sound · Computer Science 2022-03-11 Weidong Chen , Xiaofen Xing , Xiangmin Xu , Jianxin Pang , Lan Du

In this paper we propose a robust loudspeaker beamforming algorithm which is used to enhance the performance of voice driven applications in scenarios where the loudspeakers introduce the majority of the noise, e.g. when music is playing…

Audio and Speech Processing · Electrical Eng. & Systems 2025-05-21 Dimme de Groot , Baturalp Karslioglu , Odette Scharenborg , Jorge Martinez

This study introduces RUMAA, a transformer-based framework for music performance analysis that unifies score-to-performance alignment, score-informed transcription, and mistake detection in a near end-to-end manner. Unlike prior methods…

Sound · Computer Science 2025-07-17 Sungkyun Chang , Simon Dixon , Emmanouil Benetos

Transformers have achieved promising results on a variety of tasks. However, the quadratic complexity in self-attention computation has limited the applications, especially in low-resource settings and mobile or edge devices. Existing works…

Sound · Computer Science 2024-01-09 Wentao Zhu

Music generated by deep learning methods often suffers from a lack of coherence and long-term organization. Yet, multi-scale hierarchical structure is a distinctive feature of music signals. To leverage this information, we propose a…

Sound · Computer Science 2024-02-29 Manvi Agarwal , Changhong Wang , Gaël Richard

Estimating probability density and its score from samples remains a core problem in generative modeling, Bayesian inference, and kinetic theory. Existing methods are bifurcated: classical kernel density estimators (KDE) generalize across…

Machine Learning · Computer Science 2026-05-29 Vasily Ilin , Peter Sushko , Ranjay Krishna

Recent advances in Entity Resolution (ER) have leveraged Large Language Models (LLMs), achieving strong performance but at the cost of substantial computational resources or high financial overhead. Existing LLM-based ER approaches operate…

Databases · Computer Science 2026-02-06 Alexandros Zeakis , George Papadakis , Dimitrios Skoutas , Manolis Koubarakis

Optical Music Recognition (OMR) has made significant progress since its inception, with various approaches now capable of accurately transcribing music scores into digital formats. Despite these advancements, most so-called end-to-end OMR…

Computer Vision and Pattern Recognition · Computer Science 2025-06-30 Antonio Ríos-Vila , Jorge Calvo-Zaragoza , David Rizo , Thierry Paquet

The Transformer architecture has demonstrated a superior ability compared to recurrent neural networks in many different natural language processing applications. Therefore, our study applies a modified Transformer in a speech enhancement…

Audio and Speech Processing · Electrical Eng. & Systems 2021-03-04 Szu-Wei Fu , Chien-Feng Liao , Tsun-An Hsieh , Kuo-Hsuan Hung , Syu-Siang Wang , Cheng Yu , Heng-Cheng Kuo , Ryandhimas E. Zezario , You-Jin Li , Shang-Yi Chuang , Yen-Ju Lu , Yu Tsao

The goal of the audio-visual segmentation (AVS) task is to segment the sounding objects in the video frames using audio cues. However, current fusion-based methods have the performance limitations due to the small receptive field of…

Sound · Computer Science 2023-07-26 Jinxiang Liu , Chen Ju , Chaofan Ma , Yanfeng Wang , Yu Wang , Ya Zhang

AI-based music generation has made significant progress in recent years. However, generating symbolic music that is both long-structured and expressive remains a significant challenge. In this paper, we propose PerceiverS (Segmentation and…

Artificial Intelligence · Computer Science 2025-09-23 Yungang Yi , Weihua Li , Matthew Kuo , Quan Bai