English
Related papers

Related papers: GCI detection from raw speech using a fully-convol…

200 papers

In hyperspectral image (HSI) classification, spatial context has demonstrated its significance in achieving promising performance. However, conventional spatial context-based methods simply assume that spatially neighboring pixels should…

Machine Learning · Computer Science 2019-09-27 Sheng Wan , Chen Gong , Ping Zhong , Shirui Pan , Guangyu Li , Jian Yang

We propose an objective intelligibility measure (OIM), called the Gammachirp Envelope Similarity Index (GESI), that can predict speech intelligibility (SI) in older adults. GESI is a bottom-up model based on psychoacoustic knowledge from…

Audio and Speech Processing · Electrical Eng. & Systems 2025-10-30 Ayako Yamamoto , Fuki Miyazaki , Toshio Irino

This paper presents an innovative approach called BGTAI to simplify multimodal understanding by utilizing gloss-based annotation as an intermediate step in aligning Text and Audio with Images. While the dynamic temporal factors in textual…

Computer Vision and Pattern Recognition · Computer Science 2024-10-15 Sen Fang , Sizhou Chen , Yalin Feng , Xiaofeng Zhang , Teik Toe Teoh

Gravitational wave detectors now under construction are sensitive to the phase of the incident gravitational waves. Correspondingly, the signals from the different detectors can be combined, in the analysis, to simulate a single detector of…

General Relativity and Quantum Cosmology · Physics 2009-12-31 Lee Samuel Finn

Several recent work on speech synthesis have employed generative adversarial networks (GANs) to produce raw waveforms. Although such methods improve the sampling efficiency and memory usage, their sample quality has not yet reached that of…

Sound · Computer Science 2020-10-26 Jungil Kong , Jaehyeon Kim , Jaekyoung Bae

The collection of individually resolvable gravitational wave (GW) events makes up a tiny fraction of all GW signals which reach our detectors, while most lie below the confusion limit and go undetected. Like voices in a crowded room, the…

General Relativity and Quantum Cosmology · Physics 2022-03-02 Arianna I. Renzini , Boris Goncharov , Alexander C. Jenkins , Pat M. Meyers

The electroencephalography (EEG) signals recorded in parallel with speech are used to perform isolated and continuous speech recognition. During speaking process, one also hears his or her own speech and this speech perception is also…

Audio and Speech Processing · Electrical Eng. & Systems 2020-06-03 Gautam Krishna , Co Tran , Mason Carnahan , Ahmed Tewfik

Explainable AI (XAI) underwent a recent surge in research on concept extraction, focusing on extracting human-interpretable concepts from Deep Neural Networks. An important challenge facing concept extraction approaches is the difficulty of…

Machine Learning · Computer Science 2023-02-13 Dmitry Kazhdan , Botty Dimanov , Lucie Charlotte Magister , Pietro Barbiero , Mateja Jamnik , Pietro Lio

Language Identification, being an important aspect of Automatic Speaker Recognition has had many changes and new approaches to ameliorate performance over the last decade. We compare the performance of using audio spectrum in the log scale…

Computation and Language · Computer Science 2017-05-19 Vrishabh Ajay Lakhani , Rohan Mahadev

We present an exploratory framework to test whether noise-like input can induce structured responses in language models. Instead of assuming that extraterrestrial signals must be decoded, we evaluate whether inputs can trigger linguistic…

Instrumentation and Methods for Astrophysics · Physics 2025-06-04 Po-Chieh Yu

Prediction of epilepsy based on electroencephalogram (EEG) signals is a rapidly evolving field. Previous studies have traditionally applied 1D processing to the entire EEG signal. However, we have adopted the Gram Matrix method to transform…

Machine Learning · Computer Science 2025-12-16 Bihao You , Jiping Cui

Identifying speakers of quotations in narratives is an important task in literary analysis, with challenging scenarios including the out-of-domain inference for unseen speakers, and non-explicit cases where there are no speaker mentions in…

Computation and Language · Computer Science 2024-02-20 Zhenlin Su , Liyan Xu , Jin Xu , Jiangnan Li , Mingdu Huangfu

This paper proposes a framework for modeling sound change that combines deep learning and iterative learning. Acquisition and transmission of speech is modeled by training generations of Generative Adversarial Networks (GANs) on unannotated…

Computation and Language · Computer Science 2021-09-23 Gašper Beguš

Speech produced by human vocal apparatus conveys substantial non-semantic information including the gender of the speaker, voice quality, affective state, abnormalities in the vocal apparatus etc. Such information is attributed to the…

Audio and Speech Processing · Electrical Eng. & Systems 2022-08-16 Prathosh A. P. , Varun Srivastava , Mayank Mishra

This paper advances the design of CTC-based all-neural (or end-to-end) speech recognizers. We propose a novel symbol inventory, and a novel iterated-CTC method in which a second system is used to transform a noisy initial output into a…

Computation and Language · Computer Science 2022-02-24 G. Zweig , C. Yu , J. Droppo , A. Stolcke

This paper introduces a novel method to separate noisy speech into low or high frequency frames, in order to improve fundamental frequency (F0) estimation accuracy. In this proposal, the target signal is analyzed by means of the ensemble…

Audio and Speech Processing · Electrical Eng. & Systems 2021-12-21 A. Queiroz , R. Coelho

In this study, we propose a new concept, the gammachirp envelope distortion index (GEDI), based on the signal-to-distortion ratio in the auditory envelope, SDRenv to predict the intelligibility of speech enhanced by nonlinear algorithms.…

Sound · Computer Science 2020-07-21 Katsuhiko Yamamoto , Toshio Irino , Shoko Araki , Keisuke Kinoshita , Tomohiro Nakatani

For monaural speech enhancement, contextual information is important for accurate speech estimation. However, commonly used convolution neural networks (CNNs) are weak in capturing temporal contexts since they only build blocks that process…

Audio and Speech Processing · Electrical Eng. & Systems 2022-07-01 Xinmeng Xu , Yang Wang , Jie Jia , Binbin Chen , Jianjun Hao

Speech emotion recognition is a challenging task for three main reasons: 1) human emotion is abstract, which means it is hard to distinguish; 2) in general, human emotion can only be detected in some specific moments during a long…

Sound · Computer Science 2019-05-03 Yuanyuan Zhang , Jun Du , Zirui Wang , Jianshu Zhang

Audio tagging aims to assign predefined tags to audio clips to indicate the class information of audio events. Sequential audio tagging (SAT) means detecting both the class information of audio events, and the order in which they occur…

Sound · Computer Science 2022-10-25 Yuanbo Hou , Yun Wang , Wenwu Wang , Dick Botteldooren