中文
相关论文

相关论文: Exploring Pre-trained General-purpose Audio Repres…

200 篇论文

Deep learning (DL) based predictive models from electronic health records (EHR) deliver impressive performance in many clinical tasks. Large training cohorts, however, are often required to achieve high accuracy, hindering the adoption of…

计算与语言 · 计算机科学 2020-05-27 Laila Rasmy , Yang Xiang , Ziqian Xie , Cui Tao , Degui Zhi

Medical image registration is an essential topic in medical image analysis. In this paper, we propose a method for medical image registration using a pretrained large language model. We find that using the pretrained large language model to…

计算机视觉与模式识别 · 计算机科学 2024-05-30 Mingrui Ma , Yu Yang

Three-dimensional ultrasound enables real-time volumetric visualization of anatomical structures. Unlike traditional 2D ultrasound, 3D imaging reduces reliance on precise probe orientation, potentially making ultrasound more accessible to…

图像与视频处理 · 电气工程与系统科学 2026-05-05 Tristan S. W. Stevens , Oisín Nolan , Oudom Somphone , Jean-Luc Robert , Ruud J. G. van Sloun

We introduce a new audio processing technique that increases the sampling rate of signals such as speech or music using deep convolutional neural networks. Our model is trained on pairs of low and high-quality audio examples; at test-time,…

声音 · 计算机科学 2017-08-03 Volodymyr Kuleshov , S. Zayd Enam , Stefano Ermon

Musicians and audio engineers sculpt and transform their sounds by connecting multiple processors, forming an audio processing graph. However, most deep-learning methods overlook this real-world practice and assume fixed graph settings. To…

声音 · 计算机科学 2023-05-09 Sungho Lee , Jaehyun Park , Seungryeol Paik , Kyogu Lee

With the advances in deep learning, the performance of end-to-end (E2E) single-task models for speech and audio processing has been constantly improving. However, it is still challenging to build a general-purpose model with high…

音频与语音处理 · 电气工程与系统科学 2025-02-21 Xiaoyu Yang , Qiujia Li , Chao Zhang , Phil Woodland

Feature representations derived from models pre-trained on large-scale datasets have shown their generalizability on a variety of audio analysis tasks. Despite this generalizability, however, task-specific features can outperform if…

音频与语音处理 · 电气工程与系统科学 2022-06-13 Yun-Ning Hung , Alexander Lerch

While self-supervised learning (SSL) algorithms have been widely used to pre-train deep models, few efforts [11] have been done to improve representation learning of X-ray image analysis with SSL pre-trained models. In this work, we study a…

计算机视觉与模式识别 · 计算机科学 2023-10-04 Weibin Liao , Haoyi Xiong , Qingzhong Wang , Yan Mo , Xuhong Li , Yi Liu , Zeyu Chen , Siyu Huang , Dejing Dou

Digital breast tomosynthesis is rapidly replacing digital mammography as the basic x-ray technique for evaluation of the breasts. However, the sparse sampling and limited angular range gives rise to different artifacts, which manufacturers…

After constructing a deep neural network for urban sound classification, this work focuses on the sensitive application of assisting drivers suffering from hearing loss. As such, clear etiology justifying and interpreting model predictions…

声音 · 计算机科学 2021-11-22 Marco Colussi , Stavros Ntalampiras

Generalisation -- the ability of a model to perform well on unseen data -- is crucial for building reliable deepfake detectors. However, recent studies have shown that the current audio deepfake models fall short of this desideratum. In…

音频与语音处理 · 电气工程与系统科学 2024-06-14 Octavian Pascu , Adriana Stan , Dan Oneata , Elisabeta Oneata , Horia Cucu

This paper presents a paradigm that adapts general large-scale pretrained models (PTMs) to speech emotion recognition task. Although PTMs shed new light on artificial general intelligence, they are constructed with general tasks in mind,…

声音 · 计算机科学 2024-04-19 Weidong Chen , Xiaofen Xing , Peihao Chen , Xiangmin Xu

Distributed Acoustic Sensing (DAS) technology finds growing applications across various domains. However, data distribution disparities due to heterogeneous sensing environments pose challenges for data-driven artificial intelligence (AI)…

计算机视觉与模式识别 · 计算机科学 2025-08-07 Kun Gui , Hongliang Ren , Shang Shi , Jin Lu , Changqiu Yu , Quanjun Cao , Guomin Gu , Qi Xuan

While deep learning has enabled great advances in many areas of music, labeled music datasets remain especially hard, expensive, and time-consuming to create. In this work, we introduce SimCLR to the music domain and contribute a large…

声音 · 计算机科学 2021-09-28 Janne Spijkervet , John Ashley Burgoyne

Personalized virtual heart models have demonstrated increasing potential for clinical use, although the estimation of their parameters given patient-specific data remain a challenge. Traditional physics-based modeling approaches are…

信号处理 · 电气工程与系统科学 2024-03-26 Xiajun Jiang , Sumeet Vadhavkar , Yubo Ye , Maryam Toloubidokhti , Ryan Missel , Linwei Wang

Music annotation has always been one of the critical topics in the field of Music Information Retrieval (MIR). Traditional models use supervised learning for music annotation tasks. However, as supervised machine learning approaches…

音频与语音处理 · 电气工程与系统科学 2021-02-02 Yilun Zhao , Jia Guo

In this paper, we empirically investigate the effect of audio preprocessing on music tagging with deep neural networks. We perform comprehensive experiments involving audio preprocessing using different time-frequency representations,…

声音 · 计算机科学 2021-02-23 Keunwoo Choi , György Fazekas , Kyunghyun Cho , Mark Sandler

Electrocardiography analysis is widely used in various clinical applications and Deep Learning models for classification tasks are currently in the focus of research. Due to their data-driven character, they bear the potential to handle…

信号处理 · 电气工程与系统科学 2023-07-04 Theresa Bender , Philip Gemke , Ennio Idrobo-Avila , Henning Dathe , Dagmar Krefting , Nicolai Spicher

Machine hearing or listening represents an emerging area. Conventional approaches rely on the design of handcrafted features specialized to a specific audio task and that can hardly generalized to other audio fields. For example,…

计算机视觉与模式识别 · 计算机科学 2018-12-13 Imad Rida , Romain Hérault , Gilles Gasso

In this paper, we demonstrate a unique recipe to enhance the effectiveness of audio machine learning approaches by fusing pre-processing techniques into a deep learning model. Our solution accelerates training and inference performance by…

声音 · 计算机科学 2022-08-22 Devesh Khandelwal , Sean Campos , Shwetha Nagaraj , Fred Nugen , Alberto Todeschini
‹ 上一页 1 8 9 10 下一页 ›