中文
相关论文

相关论文: Accessible Fine-grained Data Representation via Sp…

200 篇论文

Machine hearing is an emerging area. Motivated by the need of a principled framework across domain applications for machine listening, we propose a generic and data-driven representation learning approach. For this sake, a novel and…

声音 · 计算机科学 2021-01-01 Imad Rida

A key aspect of machine learning models lies in their ability to learn efficient intermediate features. However, the input representation plays a crucial role in this process, and polyphonic musical scores remain a particularly complex type…

机器学习 · 计算机科学 2021-09-09 Mathieu Prang , Philippe Esling

Metric learning projects samples into an embedded space, where similarities and dissimilarities are quantified based on their learned representations. However, existing methods often rely on label-guided representation learning, where…

声音 · 计算机科学 2025-01-17 Donghuo Zeng , Kazushi Ikeda

In multimedia applications such as films and video games, spatial audio techniques are widely employed to enhance user experiences by simulating 3D sound: transforming mono audio into binaural formats. However, this process is often complex…

多媒体 · 计算机科学 2025-02-14 Xiaojing Liu , Ogulcan Gurelli , Yan Wang , Joshua Reiss

Many applications of speech technology require more and more audio data. Automatic assessment of the quality of the collected recordings is important to ensure they meet the requirements of the related applications. However, effective and…

音频与语音处理 · 电气工程与系统科学 2020-05-19 Qiang Huang , Thomas Hain

This paper investigates new data exploration experiences that enable blind users to interact with statistical data visualizations$-$bar plots, heat maps, box plots, and scatter plots$-$leveraging multimodal data representations. In addition…

人机交互 · 计算机科学 2024-03-04 JooYoung Seo , Yilin Xia , Bongshin Lee , Sean McCurry , Yu Jun Yam

Complex-valued sparse coding is a data representation which employs a dictionary of two-dimensional subspaces, while imposing a sparse, factorial prior on complex amplitudes. When trained on a dataset of natural image patches, it learns…

机器学习 · 计算机科学 2014-02-19 Wiktor Mlynarski

In music and speech, meaning is derived at multiple levels of context. Affect, for example, can be inferred both by a short sound token and by sonic patterns over a longer temporal window such as an entire recording. In this letter, we…

声音 · 计算机科学 2022-09-12 Camille Noufi , Prateek Verma

Psychoacoustical so-called "timbre spaces" map perceptual similarity ratings of instrument sounds onto low-dimensional embeddings via multidimensional scaling, but suffer from scalability issues and are incapable of generalization. Recent…

声音 · 计算机科学 2025-07-11 Haokun Tian , Stefan Lattner , Charalampos Saitis

We have recently seen great progress in learning interpretable music representations, ranging from basic factors, such as pitch and timbre, to high-level concepts, such as chord and texture. However, most methods rely heavily on music…

机器学习 · 计算机科学 2024-02-12 Xuanjie Liu , Daniel Chin , Yichen Huang , Gus Xia

Advances in deep learning have enabled the widespread deployment of speaker recognition systems (SRSs), yet they remain vulnerable to score-based impersonation attacks. Existing attacks that operate directly on raw waveforms require a large…

密码学与安全 · 计算机科学 2026-03-04 Chanwoo Hwang , Sunpill Kim , Yong Kiam Tan , Tianchi Liu , Seunghun Paik , Dongsoo Kim , Mondal Soumik , Khin Mi Mi Aung , Jae Hong Seo

In this paper, we investigate how to learn rich and robust feature representations for audio classification from visual data and acoustic images, a novel audio data modality. Former models learn audio representations from raw signals or…

计算机视觉与模式识别 · 计算机科学 2020-02-12 Andrés F. Pérez , Valentina Sanguineti , Pietro Morerio , Vittorio Murino

Multimodal automatic speech recognition systems integrate information from images to improve speech recognition quality, by grounding the speech in the visual context. While visual signals have been shown to be useful for recovering…

计算与语言 · 计算机科学 2020-10-07 Tejas Srinivasan , Ramon Sanabria , Florian Metze , Desmond Elliott

Humans can robustly recognize and localize objects by integrating visual and auditory cues. While machines are able to do the same now with images, less work has been done with sounds. This work develops an approach for dense semantic…

计算机视觉与模式识别 · 计算机科学 2020-03-10 Arun Balajee Vasudevan , Dengxin Dai , Luc Van Gool

Personalized head-related transfer functions (HRTFs) are essential for ensuring a realistic auditory experience over headphones, because they take into account individual anatomical differences that affect listening. Most machine learning…

音频与语音处理 · 电气工程与系统科学 2026-01-27 You Zhang , Andrew Francl , Ruohan Gao , Paul Calamia , Zhiyao Duan , Ishwarya Ananthabhotla

Sonification, or conveying data using non-verbal audio, is a relatively niche but growing approach for presenting data across multiple specialist domains including astronomy, climate science, and beyond. The STRAUSS Python package aims to…

天体物理仪器与方法 · 物理学 2025-05-13 James W. Trayford , Samantha Youles , Chris Harrison , Rose Shepherd , Nicolas Bonne

Low resolution fine-grained classification has widespread applicability for applications where data is captured at a distance such as surveillance and mobile photography. While fine-grained classification with high resolution images has…

计算机视觉与模式识别 · 计算机科学 2021-05-04 Maneet Singh , Shruti Nagpal , Mayank Vatsa , Richa Singh

Semantic segmentation is a challenging vision problem that usually necessitates the collection of large amounts of finely annotated data, which is often quite expensive to obtain. Coarsely annotated data provides an interesting alternative…

计算机视觉与模式识别 · 计算机科学 2018-08-03 Isay Katsman , Rohun Tripathi , Andreas Veit , Serge Belongie

Efficient audio quality assessment is vital for streamlining audio codec development. Objective assessment tools have been developed over time to algorithmically predict quality ratings from subjective assessments, the gold standard for…

音频与语音处理 · 电气工程与系统科学 2024-11-28 Pablo M. Delgado , Jürgen Herre

Even when actual technologies present the potential to augment inclusion and the United Nations has been stablished the digital access to information as a human right, people with disabilities continuously faced barriers in their…

天体物理仪器与方法 · 物理学 2023-02-02 Johanna Casado , Beatriz García , Poshak Gandhi , Wanda Díaz-Merced