中文
相关论文

相关论文: Multi-View Spectrogram Transformer for Respiratory…

200 篇论文

Medical image analysis using computer-based algorithms has attracted considerable attention from the research community and achieved tremendous progress in the last decade. With recent advances in computing resources and availability of…

图像与视频处理 · 电气工程与系统科学 2023-10-03 Huyen Tran , Duc Thanh Nguyen , John Yearwood

In this paper, we propose a novel multi-view stereo (MVS) framework that gets rid of the depth range prior. Unlike recent prior-free MVS methods that work in a pair-wise manner, our method simultaneously considers all the source images.…

计算机视觉与模式识别 · 计算机科学 2024-12-10 Yitong Dong , Yijin Li , Zhaoyang Huang , Weikang Bian , Jingbo Liu , Hujun Bao , Zhaopeng Cui , Hongsheng Li , Guofeng Zhang

Signal reconstruction from its mel-spectrogram is known as mel-spectrogram inversion and has many applications, including speech and foley sound synthesis. In this paper, we propose a mel-spectrogram inversion method based on a rigorous…

音频与语音处理 · 电气工程与系统科学 2025-01-14 Yoshiki Masuyama , Natsuki Ueno , Nobutaka Ono

Developing a reliable sound detection and recognition system offers many benefits and has many useful applications in different industries. This paper examines the difficulties that exist when attempting to perform sound classification as…

音频与语音处理 · 电气工程与系统科学 2020-08-10 Chelsea Villanueva , Joshua Vincent , Alexander Slowinski , Mohammad-Parsa Hosseini

Voice disorders significantly impact patient quality of life, yet non-invasive automated diagnosis remains under-explored due to both the scarcity of pathological voice data, and the variability in recording sources. This work introduces…

The presence of MGMT promoter methylation significantly affects how well chemotherapy works for patients with Glioblastoma Multiforme (GBM). Currently, confirmation of MGMT promoter methylation relies on invasive brain tumor tissue…

计算机视觉与模式识别 · 计算机科学 2025-12-17 Rawan Alyahya , Asrar Alruwayqi , Atheer Alqarni , Asma Alkhaldi , Metab Alkubeyyer , Xin Gao , Mona Alshahrani

The synthesis of sound via deep learning methods has recently received much attention. Some problems for deep learning approaches to sound synthesis relate to the amount of data needed to specify an audio signal and the necessity of…

声音 · 计算机科学 2022-01-10 Anastasia Natsiou , Sean O'Leary

Multi-task approaches to joint depth and segmentation prediction are well-studied for monocular images. Yet, predictions from a single-view are inherently limited, while multiple views are available in many robotics applications. On the…

计算机视觉与模式识别 · 计算机科学 2023-11-02 Mykhailo Shvets , Dongxu Zhao , Marc Niethammer , Roni Sengupta , Alexander C. Berg

Transformer has been widely used for self-supervised pre-training in Natural Language Processing (NLP) and achieved great success. However, it has not been fully explored in visual self-supervised learning. Meanwhile, previous methods only…

计算机视觉与模式识别 · 计算机科学 2021-10-26 Zhaowen Li , Zhiyang Chen , Fan Yang , Wei Li , Yousong Zhu , Chaoyang Zhao , Rui Deng , Liwei Wu , Rui Zhao , Ming Tang , Jinqiao Wang

High-dimensional structural MRI (sMRI) images are widely used for Alzheimer's Disease (AD) diagnosis. Most existing methods for sMRI representation learning rely on 3D architectures (e.g., 3D CNNs), slice-wise feature extraction with late…

计算机视觉与模式识别 · 计算机科学 2026-01-30 Dexuan Ding , Ciyuan Peng , Endrowednes Kuantama , Jingcai Guo , Jia Wu , Jian Yang , Amin Beheshti , Ming-Hsuan Yang , Yuankai Qi

Accurate biomedical image classification under low-resource conditions remains challenging due to limited annotations, subtle inter-class visual differences, and complex disease semantics. While vision--language models offer a promising…

计算机视觉与模式识别 · 计算机科学 2026-04-28 Xiaoliu Luo , Minxue Xiao , Ting Xie , Mengzhu Wang , Huiqing Qi , Joey Tianyi Zhou , Taiping Zhang , Xu Wang

In recent text-to-speech synthesis and voice conversion systems, a mel-spectrogram is commonly applied as an intermediate representation, and the necessity for a mel-spectrogram vocoder is increasing. A mel-spectrogram vocoder must solve…

声音 · 计算机科学 2022-03-07 Takuhiro Kaneko , Kou Tanaka , Hirokazu Kameoka , Shogo Seki

Natural Language Processing has recently made understanding human interaction easier, leading to improved sentimental analysis and behaviour prediction. However, the choice of words and vocal cues in conversations presents an underexplored…

计算机与社会 · 计算机科学 2022-06-24 Amna Anwar , Eiman Kanjo , Dario Ortega Anderez

In this paper, we propose a method to improve the accuracy of speech emotion recognition (SER) by using vision transformer (ViT) to attend to the correlation of frequency (y-axis) with time (x-axis) in spectrogram and transferring…

声音 · 计算机科学 2024-11-05 Jeong-Yoon Kim , Seung-Ho Lee

Pattern recognition from audio signals is an active research topic encompassing audio tagging, acoustic scene classification, music classification, and other areas. Spectrogram and mel-frequency cepstral coefficients (MFCC) are among the…

音频与语音处理 · 电气工程与系统科学 2022-11-18 Md. Istiaq Ansari , Taufiq Hasan

An embedding-based speaker adaptive training (SAT) approach is proposed and investigated in this paper for deep neural network acoustic modeling. In this approach, speaker embedding vectors, which are a constant given a particular speaker,…

计算与语言 · 计算机科学 2017-10-20 Xiaodong Cui , Vaibhava Goel , George Saon

Deep neural networks have been applied to audio spectrograms for respiratory sound classification, but it remains challenging to achieve satisfactory performance due to the scarcity of available data. Moreover, domain mismatch may be…

音频与语音处理 · 电气工程与系统科学 2025-06-16 Peidong Wei , Shiyu Miao , Lin Li

Spectrograms visualize the frequency components of a given signal which may be an audio signal or even a time-series signal. Audio signals have higher sampling rate and high variability of frequency with time. Spectrograms can capture such…

信号处理 · 电气工程与系统科学 2021-09-06 Sidharth Srivatsav Sribhashyam , Md Sirajus Salekin , Dmitry Goldgof , Ghada Zamzmi , Mark Last , Yu Sun

The two-dimensional (2D) numerical approaches for vocal tract (VT) modelling can afford a better balance between the low computational cost and accurate rendering of acoustic wave propagation. However, they require a high spatio-temporal…

声音 · 计算机科学 2021-02-10 Debasish Ray Mohapatra , Victor Zappi , Sidney Fels

Active research is currently underway to enhance the efficiency of vision transformers (ViTs). Most studies have focused solely on effective token mixers, overlooking the potential relationship with normalization. To boost diverse feature…

计算机视觉与模式识别 · 计算机科学 2024-12-02 Jongseong Bae , Susang Kim , Minsu Cho , Ha Young Kim