中文
相关论文

相关论文: Trigonometric dictionary based codec for music com…

200 篇论文

A key aspect of machine learning models lies in their ability to learn efficient intermediate features. However, the input representation plays a crucial role in this process, and polyphonic musical scores remain a particularly complex type…

机器学习 · 计算机科学 2021-09-09 Mathieu Prang , Philippe Esling

Natural signals and images are well-known to be approximately sparse in transform domains such as Wavelets and DCT. This property has been heavily exploited in various applications in image processing and medical imaging. Compressed sensing…

机器学习 · 计算机科学 2015-10-26 Saiprasad Ravishankar , Yoram Bresler

Music auto-tagging is essential for organizing and discovering music in extensive digital libraries. While foundation models achieve exceptional performance in this domain, their outputs often lack interpretability, limiting trust and…

机器学习 · 计算机科学 2026-05-28 Andreas Patakis , Vassilis Lyberatos , Spyridon Kantarelis , Edmund Dervakos , Giorgos Stamou

This thesis develops a Transformer model based on Whisper, which extracts melodies and chords from music audio and records them into ABC notation. A comprehensive data processing workflow is customized for ABC notation, including data…

声音 · 计算机科学 2024-10-23 Hongyao Zhang , Bohang Sun

EEG and audio are inherently distinct modalities, differing in sampling rate, channel structure, and scale. Yet, we show that pretrained neural audio codecs can serve as effective starting points for EEG compression, provided that the data…

机器学习 · 计算机科学 2025-12-01 Ard Kastrati , Luca Lanzendörfer , Riccardo Rigoni , John Staib Matilla , Roger Wattenhofer

An ever increasing amount of our digital communication, media consumption, and content creation revolves around videos. We share, watch, and archive many aspects of our lives through them, all of which are powered by strong video…

计算机视觉与模式识别 · 计算机科学 2018-04-20 Chao-Yuan Wu , Nayan Singhal , Philipp Krähenbühl

Recent video codecs with multiple separable transforms can achieve significant coding gains using asymmetric trigonometric transforms (DCTs and DSTs), because they can exploit diverse statistics of residual block signals. However, they add…

图像与视频处理 · 电气工程与系统科学 2025-05-30 Amir Said , Hilmi E. Egilmez , Yung-Hsuan Chao

Existing convex relaxation-based approaches to reconstruction in compressed sensing assume that noise in the measurements is independent of the signal of interest. We consider the case of noise being linearly correlated with the signal and…

信息论 · 计算机科学 2014-01-03 Thomas Arildsen , Torben Larsen

The standard approach to compressive sampling considers recovering an unknown deterministic signal with certain known structure, and designing the sub-sampling pattern and recovery algorithm based on the known structure. This approach…

信息论 · 计算机科学 2016-02-03 Yen-Huan Li , Volkan Cevher

The MUSIC algorithm, with its extension for imaging sparse {\em extended} objects, is analyzed by compressed sensing (CS) techniques. The notion of restricted isometry property (RIP) and an upper bound on the restricted isometry constant…

信息论 · 计算机科学 2015-05-19 Albert C. Fannjiang

Photometric redshift surveys map the distribution of matter in the Universe through the positions and shapes of galaxies with poorly resolved measurements of their radial coordinates. While a tomographic analysis can be used to recover some…

宇宙学与河外天体物理 · 物理学 2017-12-06 David Alonso

For the lossless compression of dynamic 3-D+t volumes as produced by medical devices like Computed Tomography, various coding schemes can be applied. This paper shows that 3-D subband coding outperforms lossless HEVC coding and additionally…

图像与视频处理 · 电气工程与系统科学 2023-02-03 Daniela Lanz , Jürgen Seiler , Karina Jaskolka , André Kaup

This paper introduces effective design choices for text-to-music retrieval systems. An ideal text-based retrieval system would support various input queries such as pre-defined tags, unseen tags, and sentence-level descriptions. In reality,…

信息检索 · 计算机科学 2022-11-29 SeungHeon Doh , Minz Won , Keunwoo Choi , Juhan Nam

We present a framework based on neural networks to extract music scores directly from polyphonic audio in an end-to-end fashion. Most previous Automatic Music Transcription (AMT) methods seek a piano-roll representation of the pitches, that…

声音 · 计算机科学 2019-10-29 Miguel A. Román , Antonio Pertusa , Jorge Calvo-Zaragoza

Compressing the sign information of discrete cosine transform (DCT) coefficients is an intractable problem in image coding schemes due to the equiprobable characteristics of the signs. To overcome this difficulty, we propose an efficient…

信息论 · 计算机科学 2024-05-13 Kei Suzuki , Chihiro Tsutake , Keita Takahashi , Toshiaki Fujii

Optical Music Recognition (OMR) aims to convert printed or handwritten music score images into editable symbolic representations. This paper presents an end-to-end OMR framework that combines residual bottleneck convolutions with…

计算机视觉与模式识别 · 计算机科学 2026-04-21 Junwen Ma , Huhu Xue , Xingyuan Zhao , and Weicheng Fu

We proposed a practical ECG compression system which is beneficial for tele-monitoring cardiovascular diseases. There are two steps in the compression framework. First, we partition ECG signal into segments according to R- to R-wave…

信息论 · 计算机科学 2017-04-19 Pengda Wong

Nowadays, real-time video communication over the internet through video conferencing applications has become an invaluable tool in everyone's professional and personal life. This trend underlines the need for video coding algorithms that…

多媒体 · 计算机科学 2015-10-05 Stamos Katsigiannis , Georgios Papaioannou , Dimitris Maroulis

A combinatorial approach to compressive sensing based on a deterministic column replacement technique is proposed. Informally, it takes as input a pattern matrix and ingredient measurement matrices, and results in a larger measurement…

信息论 · 计算机科学 2014-03-10 Charles J. Colbourn , Daniel Horsley , Violet R. Syrotiuk

As the parameter size of large language models (LLMs) continues to expand, the need for a large memory footprint and high communication bandwidth have become significant bottlenecks for the training and inference of LLMs. To mitigate these…

机器学习 · 计算机科学 2024-07-02 Ceyu Xu , Yongji Wu , Xinyu Yang , Beidi Chen , Matthew Lentz , Danyang Zhuo , Lisa Wu Wills