English
Related papers

Related papers: Boosting the Predictive Accurary of Singer Identif…

200 papers

Language Models pretrained on large textual data have been shown to encode different types of knowledge simultaneously. Traditionally, only the features from the last layer are used when adapting to new tasks or data. We put forward that,…

Computation and Language · Computer Science 2024-05-08 Muhammad ElNokrashy , Badr AlKhamissi , Mona Diab

Recent advances in unsupervised speech representation learning discover new approaches and provide new state-of-the-art for diverse types of speech processing tasks. This paper presents an investigation of using wav2vec 2.0 deep speech…

The challenge of image generation has been effectively modeled as a problem of structure priors or transformation. However, existing models have unsatisfactory performance in understanding the global input image structures because of…

Computer Vision and Pattern Recognition · Computer Science 2023-10-26 Pourya Shamsolmoali , Masoumeh Zareapoor , Huiyu Zhou , Xuelong Li , Yue Lu

Music source separation (MSS) aims to extract 'vocals', 'drums', 'bass' and 'other' tracks from a piece of mixed music. While deep learning methods have shown impressive results, there is a trend toward larger models. In our paper, we…

Audio and Speech Processing · Electrical Eng. & Systems 2024-03-20 Junyu Chen , Susmitha Vekkot , Pancham Shukla

Recently, the end-to-end approach that learns hierarchical representations from raw data using deep convolutional neural networks has been successfully explored in the image, text and speech domains. This approach was applied to musical…

Sound · Computer Science 2017-05-23 Jongpil Lee , Jiyoung Park , Keunhyoung Luke Kim , Juhan Nam

The Internet has turned the entire world into a small village;this is because it has made it possible to share millions of images and videos. However, sending and receiving a huge amount of data is considered to be a main challenge. To…

Computer Vision and Pattern Recognition · Computer Science 2023-03-24 Hassan Mohamed Muhi-Aldeen , Asma A. Abdulrahman , Jabbar Abed Eleiwy , Fouad S. Tahir , Yurii Khlaponin

Power measurement algorithms based on Fourier transform are susceptible to errors caused by interharmonics, while wavelet transform algorithms are particularly sensitive to even harmonics due to band decomposition effects. The empirical…

Signal Processing · Electrical Eng. & Systems 2025-02-17 Jian Liu , Wei Zhao , Shisong Li

Automated detection of voice disorders with computational methods is a recent research area in the medical domain since it requires a rigorous endoscopy for the accurate diagnosis. Efficient screening methods are required for the diagnosis…

Quantitative Methods · Quantitative Biology 2018-12-06 Vibhuti Gupta

Vocal Percussion Transcription (VPT) is concerned with the automatic detection and classification of vocal percussion sound events, allowing music creators and producers to sketch drum lines on the fly. Classifier algorithms in VPT systems…

Sound · Computer Science 2022-04-12 Alejandro Delgado , Emir Demirel , Vinod Subramanian , Charalampos Saitis , Mark Sandler

Change detection (CD) aims to detect change regions within an image pair captured at different times, playing a significant role in diverse real-world applications. Nevertheless, most of the existing works focus on designing advanced…

Computer Vision and Pattern Recognition · Computer Science 2024-12-10 Qing Guo , Ruofei Wang , Rui Huang , Shuifa Sun , Yuxiang Zhang

A central problem in machine learning and pattern recognition is the process of recognizing the most important features. In this paper, we provide a new feature selection method (DRPT) that consists of first removing the irrelevant features…

Machine Learning · Computer Science 2021-05-19 Majid Afshar , Hamid Usefi

The expressive variability in producing a musical note conveys information essential to the modeling of orchestration and style. As such, it plays a crucial role in computer-assisted browsing of massive digital music corpora. Yet, although…

Sound · Computer Science 2018-08-30 Vincent Lostanlen , Joakim Andén , Mathieu Lagrange

With the continuous improvements of deepfake methods, forgery messages have transitioned from single-modality to multi-modal fusion, posing new challenges for existing forgery detection algorithms. In this paper, we propose AVT2-DWF, the…

Computer Vision and Pattern Recognition · Computer Science 2024-03-25 Rui Wang , Dengpan Ye , Long Tang , Yunming Zhang , Jiacheng Deng

Automatic lyrics transcription (ALT) remains a challenging task in the field of music information retrieval, despite great advances in automatic speech recognition (ASR) brought about by transformer-based architectures in recent years. One…

Sound · Computer Science 2025-06-19 Jaza Syed , Ivan Meresman Higgs , Ondřej Cífka , Mark Sandler

A novel digital watermarking for ownership verification and image authentication applications using discrete wavelet transform (DWT) is proposed in this paper. Most previous proposed watermarking algorithms embed sequences of random numbers…

Cryptography and Security · Computer Science 2012-06-21 Mehdi Khalili

For most of the state-of-the-art speech enhancement techniques, a spectrogram is usually preferred than the respective time-domain raw data since it reveals more compact presentation together with conspicuous temporal information over a…

Sound · Computer Science 2016-08-24 Syu-Siang Wang , Alan Chern , Yu Tsao , Jeih-weih Hung , Xugang Lu , Ying-Hui Lai , Borching Su

Accurate classification of sleep stages is crucial for the diagnosis and management of sleep disorders. Conventional approaches for sleep scoring rely on manual annotation or features extracted from EEG signals in the time or frequency…

Machine Learning · Computer Science 2025-10-10 Mehdi Zekriyapanah Gashti , Ghasem Farjamnia

We introduce the use of DCTNet, an efficient approximation and alternative to PCANet, for acoustic signal classification. In PCANet, the eigenfunctions of the local sample covariance matrix (PCA) are used as filterbanks for convolution and…

Sound · Computer Science 2016-05-09 Yin Xian , Andrew Thompson , Xiaobai Sun , Douglas Nowacek , Loren Nolte

Modern applications such as voice recognition rely on the ability to compare signals to pre-recorded ones to classify them. However, this comparison typically needs to ignore differences due to signal noise, temporal offset, signal…

Machine Learning · Computer Science 2022-06-16 Arvind Seshan

Biomedical signal classification presents unique challenges due to long sequences, complex temporal dynamics, and multi-scale frequency patterns that are poorly captured by standard transformer architectures. We propose WaveFormer, a…

Machine Learning · Computer Science 2026-02-13 Habib Irani , Bikram De , Vangelis Metsis