中文
相关论文

相关论文: HarmoF0: Logarithmic Scale Dilated Convolution For…

200 篇论文

Broadband wireless channel is a time dispersive and becomes strongly frequency selective. In most cases, the channel is composed of a few dominant coefficients and a large part of coefficients is approximately zero or zero. To exploit the…

信息论 · 计算机科学 2010-05-14 Guan Gui , Wei Peng , Qun Wan , Fumiyuki Adachi

We investigate the task of retrieving information from compositional distributed representations formed by Hyperdimensional Computing/Vector Symbolic Architectures and present novel techniques which achieve new information rate bounds.…

The aim of latent variable disentanglement is to infer the multiple informative latent representations that lie behind a data generation process and is a key factor in controllable data generation. In this paper, we propose a deep neural…

声音 · 计算机科学 2023-09-07 Yiming Wu

Structure perception is a fundamental aspect of music cognition in humans. Historically, the hierarchical organization of music into structures served as a narrative device for conveying meaning, creating expectancy, and evoking emotions in…

声音 · 计算机科学 2023-03-28 Nicolas Lazzari , Andrea Poltronieri , Valentina Presutti

The voting method, an ensemble approach for fundamental frequency estimation, is empirically known for its robustness but lacks thorough investigation. This paper provides a principled analysis and improvement of this technique. First, we…

声音 · 计算机科学 2026-02-03 Junya Koguchi , Tomoki Koriyama

The field-of-view is an important metric when designing a model for semantic segmentation. To obtain a large field-of-view, previous approaches generally choose to rapidly downsample the resolution, usually with average poolings or stride 2…

计算机视觉与模式识别 · 计算机科学 2023-11-29 Roland Gao

We approach the singing phrase audio to score matching problem by using phonetic and duration information - with a focus on studying the jingju a cappella singing case. We argue that, due to the existence of a basic melodic contour for each…

声音 · 计算机科学 2017-07-13 Rong Gong , Jordi Pons , Xavier Serra

Modern audio source separation techniques rely on optimizing sequence model architectures such as, 1D-CNNs, on mixture recordings to generalize well to unseen mixtures. Specifically, recent focus is on time-domain based architectures such…

Multi-channel speech enhancement extracts speech using multiple microphones that capture spatial cues. Effectively utilizing directional information is key for multi-channel enhancement. Deep learning shows great potential on multi-channel…

声音 · 计算机科学 2023-09-21 Jiahui Pan , Pengjie Shen , Hui Zhang , Xueliang Zhang

Fundamental frequency is one of the most important characteristics of speech and audio signals. Harmonic model-based fundamental frequency estimators offer a higher estimation accuracy and robustness against noise than the widely used…

Coded caching provides significant gains over conventional uncoded caching by creating multicasting opportunities among distinct requests. Massive multiple-input multiple-output (MIMO) systems require downlink channel state information…

信息论 · 计算机科学 2019-07-08 Qianqian Yang , Mahdi Boloursaz Mashhadi , Deniz Gündüz

In this paper, we introduce a new framework for unsupervised deep homography estimation. Our contributions are 3 folds. First, unlike previous methods that regress 4 offsets for a homography, we propose a homography flow representation,…

计算机视觉与模式识别 · 计算机科学 2021-08-19 Nianjin Ye , Chuan Wang , Haoqiang Fan , Shuaicheng Liu

Recently, convolutional neural network (CNN) techniques have gained popularity as a tool for hyperspectral image classification (HSIC). To improve the feature extraction efficiency of HSIC under the condition of limited samples, the current…

计算机视觉与模式识别 · 计算机科学 2023-04-06 Hongmin Gao , Zhonghao Chen , Chenming Li

This paper presents a geometric approach to pitch estimation (PE)-an important problem in Music Information Retrieval (MIR), and a precursor to a variety of other problems in the field. Though there exist a number of highly-accurate…

声音 · 计算机科学 2020-12-09 Tom Goodman , Karoline van Gemst , Peter Tino

In high-dynamic range (HDR) analog-to-digital converters (ADCs), having many quantization bits minimizes quantization errors but results in high bit rates, limiting their application scope. A strategy combining modulo-folding with a low-DR…

信号处理 · 电气工程与系统科学 2023-11-23 Satish Mulleti , Resham Yashwanth Kumar , Laxmeesha Somappa

Extraction of the predominant pitch from polyphonic audio is one of the fundamental tasks in the field of music information retrieval and computational musicology. To accomplish this task using machine learning, a large amount of labeled…

音频与语音处理 · 电气工程与系统科学 2023-04-07 Kavya Ranjan Saxena , Vipul Arora

In this paper, we present an efficient neural network for end-to-end general purpose audio source separation. Specifically, the backbone structure of this convolutional network is the SUccessive DOwnsampling and Resampling of…

音频与语音处理 · 电气工程与系统科学 2021-05-14 Efthymios Tzinis , Zhepei Wang , Paris Smaragdis

Neural vocoders have recently advanced waveform generation, yielding natural and expressive audio. Among these approaches, iSTFT-based vocoders have recently gained attention. They predict a complex-valued spectrogram and then synthesize…

声音 · 计算机科学 2026-03-13 Hyung-Seok Oh , Deok-Hyeon Cho , Seung-Bin Kim , Seong-Whan Lee

Malware detection is an interesting and valuable domain to work in because it has significant real-world impact and unique machine-learning challenges. We investigate existing long-range techniques and benchmarks and find that they're not…

密码学与安全 · 计算机科学 2024-03-28 Mohammad Mahmudul Alam , Edward Raff , Stella Biderman , Tim Oates , James Holt

Balancing reconstruction quality versus model efficiency remains a critical challenge in lightweight single image super-resolution (SISR). Despite the prevalence of attention mechanisms in recent state-of-the-art SISR approaches that…

计算机视觉与模式识别 · 计算机科学 2025-05-28 M. Akin Yilmaz , Ahmet Bilican , A. Murat Tekalp