English
Related papers

Related papers: Wavelet Scattering Transform for Improving General…

200 papers

Several recent papers have published good solutions for language identification (LID) for about 300 high-resource and medium-resource languages. However, there is no LID available that (i) covers a wide range of low-resource languages, (ii)…

Computation and Language · Computer Science 2024-07-04 Amir Hossein Kargaran , Ayyoob Imani , François Yvon , Hinrich Schütze

Time-frequency representations (TFRs) of signals, such as the windowed Fourier transform (WFT), wavelet transform (WT) and their synchrosqueezed variants (SWFT, SWT), provide powerful analysis tools. However, there are many important issues…

Numerical Analysis · Mathematics 2014-05-27 Dmytro Iatsenko , Peter V. E. McClintock , Aneta Stefanovska

Artificial intelligence has achieved notable results in sign language recognition and translation. However, relatively few efforts have been made to significantly improve the quality of life for the 72 million hearing-impaired people…

Computers and Society · Computer Science 2025-01-14 Yulong Li , Bolin Ren , Ke Hu , Changyuan Liu , Zhengyong Jiang , Kang Dang , Jionglong Su

While large language models (LLMs) have been applied to automatic speech recognition (ASR), the task of making the model streamable remains a challenge. This paper proposes a novel model architecture, Transducer-Llama, that integrates LLMs…

Computation and Language · Computer Science 2024-12-24 Keqi Deng , Jinxi Guo , Yingyi Ma , Niko Moritz , Philip C. Woodland , Ozlem Kalinli , Mike Seltzer

Due to the mismatch of statistical distributions of acoustic speech between training and testing sets, the performance of spoken language identification (SLID) could be drastically degraded. In this paper, we propose an unsupervised neural…

Machine Learning · Computer Science 2020-12-25 Xugang Lu , Peng Shen , Yu Tsao , Hisashi Kawai

The disparity in phonology between learner's native (L1) and target (L2) language poses a significant challenge for mispronunciation detection and diagnosis (MDD) systems. This challenge is further intensified by lack of annotated L2 data.…

Sound · Computer Science 2023-08-08 Yassine El Kheir , Shammur Absar Chowdhury , Ahmed Ali

The dream of machine learning in materials science is for a model to learn the underlying physics of an atomic system, allowing it to move beyond interpolation of the training set to the prediction of properties that were not present in the…

Computational Physics · Physics 2020-08-26 Paul Sinz , Michael W. Swift , Xavier Brumwell , Jialin Liu , Kwang Jin Kim , Yue Qi , Matthew Hirn

Fine-tuning pretrained language models (LMs) without making any architectural changes has become a norm for learning various language downstream tasks. However, for non-language downstream tasks, a common practice is to employ task-specific…

Self-supervised learning (SSL)-based speech models are extensively used for full-stack speech processing. However, it has been observed that improving SSL-based speech representations using unlabeled speech for content-related tasks is…

Computation and Language · Computer Science 2024-06-14 Amit Meghanani , Thomas Hain

In this paper, we propose VoiceID loss, a novel loss function for training a speech enhancement model to improve the robustness of speaker verification. In contrast to the commonly used loss functions for speech enhancement such as the L2…

Audio and Speech Processing · Electrical Eng. & Systems 2019-07-08 Suwon Shon , Hao Tang , James Glass

Current theoretical treatment of mode splitting and scattering loss resulting from sub-wavelength scatterers attached to the surface of high-quality-factor whispering-gallery-mode microresonators is not satisfactory. Different models have…

Optics · Physics 2015-06-05 Qing Li , Ali A. Eftekhar , Zhixuan Xia , Ali Adibi

Variational Autoencoders (VAEs) are powerful generative models, however their generated samples are known to suffer from a characteristic blurriness, as compared to the outputs of alternative generating techniques. Extensive research…

Image and Video Processing · Electrical Eng. & Systems 2024-01-09 Vibhu Dalal

Language identification (LID) recognizes the language of a spoken utterance automatically. According to recent studies, LID models trained with an automatic speech recognition (ASR) task perform better than those trained with a LID task…

Audio and Speech Processing · Electrical Eng. & Systems 2023-04-17 Jinseok Park , Hyung Yong Kim , Jihwan Park , Byeong-Yeol Kim , Shukjae Choi , Yunkyu Lim

In recent years, deep networks have led to dramatic improvements in speech enhancement by framing it as a data-driven pattern recognition problem. In many modern enhancement systems, large amounts of data are used to train a deep network to…

Stuttering is a varied speech disorder that harms an individual's communication ability. Persons who stutter (PWS) often use speech therapy to cope with their condition. Improving speech recognition systems for people with such non-typical…

Audio and Speech Processing · Electrical Eng. & Systems 2026-01-27 Sebastian P. Bayerl , Dominik Wagner , Elmar Nöth , Korbinian Riedhammer

By combining the undecimated wavelet transform within a Word Embedded Semantic Marginal Autoencoder (WESMA), this research study provides a novel strategy for improving security measures and denoising multiple languages. The incorporation…

Computation and Language · Computer Science 2023-07-10 Shreyanth S

Generating highly detailed, complex data is a long-standing and frequently considered problem in the machine learning field. However, developing detail-aware generators remains an challenging and open problem. Generative adversarial…

Machine Learning · Computer Science 2022-09-07 Lukas Prantl , Jan Bender , Tassilo Kugelstadt , Nils Thuerey

Sign language translation (SLT) is challenging, as it involves converting sign language videos into natural language. Previous studies have prioritized accuracy over diversity. However, diversity is crucial for handling lexical and…

Computer Vision and Pattern Recognition · Computer Science 2024-11-27 JiHwan Moon , Jihoon Park , Jungeun Kim , Jongseong Bae , Hyeongwoo Jeon , Ha Young Kim

Deep learning has dramatically improved the performance of speech recognition systems through learning hierarchies of features optimized for the task at hand. However, true end-to-end learning, where features are learned directly from…

Computation and Language · Computer Science 2016-04-06 Zhenyao Zhu , Jesse H. Engel , Awni Hannun

Marine mammal communication is a complex field, hindered by the diversity of vocalizations and environmental factors. The Watkins Marine Mammal Sound Database (WMMD) constitutes a comprehensive labeled dataset employed in machine learning…

Signal Processing · Electrical Eng. & Systems 2024-06-27 Alessandro Licciardi , Davide Carbone
‹ Prev 1 3 4 5 6 7 10 Next ›