English
Related papers

Related papers: Multi-View Spectrogram Transformer for Respiratory…

200 papers

Quality and intelligibility of speech signals are degraded under additive background noise which is a critical problem for hearing aid and cochlear implant users. Motivated to address this problem, we propose a novel speech enhancement…

Audio and Speech Processing · Electrical Eng. & Systems 2019-11-06 Hamidreza Baradaran Kashani , Ata Jodeiri , Mohammad Mohsen Goodarzi , Iman Sarraf Rezaei

Purely neural network (NN) based speech separation and enhancement methods, although can achieve good objective scores, inevitably cause nonlinear speech distortions that are harmful for the automatic speech recognition (ASR). On the other…

Audio and Speech Processing · Electrical Eng. & Systems 2020-08-03 Yong Xu , Meng Yu , Shi-Xiong Zhang , Lianwu Chen , Chao Weng , Jianming Liu , Dong Yu

Traditional deep learning approaches for breast cancer classification has predominantly concentrated on single-view analysis. In clinical practice, however, radiologists concurrently examine all views within a mammography exam, leveraging…

Computer Vision and Pattern Recognition · Computer Science 2024-02-27 Sushmita Sarker , Prithul Sarker , George Bebis , Alireza Tavakkoli

We present Multiscale Vision Transformers (MViT) for video and image recognition, by connecting the seminal idea of multiscale feature hierarchies with transformer models. Multiscale Transformers have several channel-resolution scale…

Computer Vision and Pattern Recognition · Computer Science 2021-04-23 Haoqi Fan , Bo Xiong , Karttikeya Mangalam , Yanghao Li , Zhicheng Yan , Jitendra Malik , Christoph Feichtenhofer

Despite its widespread adoption as the prominent neural architecture, the Transformer has spurred several independent lines of work to address its limitations. One such approach is selective state space models, which have demonstrated…

Sound · Computer Science 2024-06-11 Sarthak Yadav , Zheng-Hua Tan

This paper presents an audio visual automatic speech recognition (AV-ASR) system using a Transformer-based architecture. We particularly focus on the scene context provided by the visual information, to ground the ASR. We extract…

Audio and Speech Processing · Electrical Eng. & Systems 2020-05-01 Georgios Paraskevopoulos , Srinivas Parthasarathy , Aparna Khare , Shiva Sundaram

The core of Multi-view Stereo(MVS) is the matching process among reference and source pixels. Cost aggregation plays a significant role in this process, while previous methods focus on handling it via CNNs. This may inherit the natural…

Computer Vision and Pattern Recognition · Computer Science 2023-05-18 Weitao Chen , Hongbin Xu , Zhipeng Zhou , Yang Liu , Baigui Sun , Wenxiong Kang , Xuansong Xie

Deep learning has dramatically improved the performance of sounds recognition. However, learning acoustic models directly from the raw waveform is still challenging. Current waveform-based models generally use time-domain convolutional…

Sound · Computer Science 2018-03-29 Boqing Zhu , Changjian Wang , Feng Liu , Jin Lei , Zengquan Lu , Yuxing Peng

Recently, MLP structures have regained popularity, with MLP-Mixer standing out as a prominent example. In the field of computer vision, MLP-Mixer is noted for its ability to extract data information from both channel and token perspectives,…

Machine Learning · Computer Science 2024-03-05 Qingfeng Ji , Yuxin Wang , Letong Sun

In short video and live broadcasts, speech, singing voice, and background music often overlap and obscure each other. This complexity creates difficulties in structuring and recognizing the audio content, which may impair subsequent ASR and…

Sound · Computer Science 2024-04-18 Ye Bai , Chenxing Li , Hao Li , Yuanyuan Zhao , Xiaorui Wang

Audio DNNs have demonstrated impressive performance on various machine listening tasks; however, most of their representations are computationally costly and uninterpretable, leaving room for optimization. Here, we propose a novel approach…

Sound · Computer Science 2025-08-20 Andrew Chang , Yike Li , Iran R. Roman , David Poeppel

In multi-speaker speech synthesis, data from a number of speakers usually tend to have great diversity due to the fact that the speakers may differ largely in ages, speaking styles, emotions, and so on. It is important but challenging to…

Sound · Computer Science 2022-02-14 Qinghua Wu , Quanbo Shen , Jian Luan , YuJun Wang

Respiratory sound classification (RSC) is challenging due to varied acoustic signatures, primarily influenced by patient demographics and recording environments. To address this issue, we introduce a text-audio multimodal model that…

Sound · Computer Science 2024-06-17 June-Woo Kim , Miika Toikkanen , Yera Choi , Seoung-Eun Moon , Ho-Young Jung

Deep learning provides an excellent avenue for optimizing diagnosis and patient monitoring for clinical-based applications, which can critically enhance the response time to the onset of various conditions. For cardiovascular disease, one…

Machine Learning · Computer Science 2023-02-23 Ankur Samanta , Mark Karlov , Meghna Ravikumar , Christian McIntosh Clarke , Jayakumar Rajadas , Kaveh Hassani

Vocoders are models capable of transforming a low-dimensional spectral representation of an audio signal, typically the mel spectrogram, to a waveform. Modern speech generation pipelines use a vocoder as their final component. Recent…

Sound · Computer Science 2022-08-29 Bruno Di Giorgi , Mark Levy , Richard Sharp

The global demand for radiologists is increasing rapidly due to a growing reliance on medical imaging services, while the supply of radiologists is not keeping pace. Advances in computer vision and image processing technologies present…

Computer Vision and Pattern Recognition · Computer Science 2025-11-11 Shehroz S. Khan , Petar Przulj , Ahmed Ashraf , Ali Abedi

Audio and video are two most common modalities in the mainstream media platforms, e.g., YouTube. To learn from multimodal videos effectively, in this work, we propose a novel audio-video recognition approach termed audio video Transformer,…

Computer Vision and Pattern Recognition · Computer Science 2024-01-10 Wentao Zhu

Due to the powerful ability in capturing the global information, Transformer has become an alternative architecture of CNNs for hyperspectral image classification. However, general Transformer mainly considers the global spectral…

Computer Vision and Pattern Recognition · Computer Science 2023-10-31 Zhiqiang Gong , Xian Zhou , Wen Yao

Single-channel speech separation in time domain and frequency domain has been widely studied for voice-driven applications over the past few years. Most of previous works assume known number of speakers in advance, however, which is not…

Audio and Speech Processing · Electrical Eng. & Systems 2020-04-02 Yiming Xiao , Haijian Zhang

Adding visual cues to audio-based speech separation can improve separation performance. This paper introduces AV-CrossNet, an audiovisual (AV) system for speech enhancement, target speaker extraction, and multi-talker speaker separation.…

Audio and Speech Processing · Electrical Eng. & Systems 2024-06-18 Vahid Ahmadi Kalkhorani , Cheng Yu , Anurag Kumar , Ke Tan , Buye Xu , DeLiang Wang
‹ Prev 1 3 4 5 6 7 10 Next ›