English
Related papers

Related papers: POLIPHONE: A Dataset for Smartphone Model Identifi…

200 papers

Smartphones have been employed with biometric-based verification systems to provide security in highly sensitive applications. Audio-visual biometrics are getting popular due to their usability, and also it will be challenging to spoof…

In this article, the authors discuss the problem of forensic authentication of digital audio recordings. Although forensic audio has been addressed in several articles, the existing approaches are focused on analog magnetic recordings,…

Cryptography and Security · Computer Science 2022-03-15 Marcos Faundez-Zanuy , Jose Juan Lucena-Molina , Martin Hagmueller

Given the large number of new musical tracks released each year, automated approaches to plagiarism detection are essential to help us track potential violations of copyright. Most current approaches to plagiarism detection are based on…

With the proliferation of speech deepfake generators, it becomes crucial not only to assess the authenticity of synthetic audio but also to trace its origin. While source attribution models attempt to address this challenge, they often…

Sound · Computer Science 2025-05-21 Viola Negroni , Davide Salvi , Paolo Bestagini , Stefano Tubaro

Storytelling is multi-modal in the real world. When one tells a story, one may use all of the visualizations and sounds along with the story itself. However, prior studies on storytelling datasets and tasks have paid little attention to…

Multimedia · Computer Science 2023-10-31 Jaeyeon Bae , Seokhoon Jeong , Seokun Kang , Namgi Han , Jae-Yon Lee , Hyounghun Kim , Taehwan Kim

The rapid progress of deep speech synthesis models has posed significant threats to society such as malicious manipulation of content. This has led to an increase in studies aimed at detecting so-called deepfake audio. However, existing…

Sound · Computer Science 2024-11-19 Xinrui Yan , Jiangyan Yi , Jianhua Tao , Jie Chen

Beginner musicians often struggle to identify specific errors in their performances, such as playing incorrect notes or rhythms. There are two limitations in existing tools for music error detection: (1) Existing approaches rely on…

Generative models are now capable of synthesizing images, speeches, and videos that are hardly distinguishable from authentic contents. Such capabilities cause concerns such as malicious impersonation and IP theft. This paper investigates a…

Sound · Computer Science 2022-03-16 Yongbaek Cho , Changhoon Kim , Yezhou Yang , Yi Ren

Machine learning techniques have proved useful for classifying and analyzing audio content. However, recent methods typically rely on abstract and high-dimensional representations that are difficult to interpret. Inspired by…

Audio is a rich sensing modality that is useful for a variety of human activity recognition tasks. However, the ubiquitous nature of smartphones and smart speakers with always-on microphones has led to numerous privacy concerns and a lack…

We consider the problem of audio voice separation for binaural applications, such as earphones and hearing aids. While today's neural networks perform remarkably well (separating $4+$ sources with 2 microphones) they assume a known or fixed…

Sound · Computer Science 2022-07-18 Zhongweiyang Xu , Romit Roy Choudhury

We introduce PodcastMix, a dataset formalizing the task of separating background music and foreground speech in podcasts. We aim at defining a benchmark suitable for training and evaluating (deep learning) source separation models. To that…

Sound · Computer Science 2022-07-18 Nicolás Schmidt , Jordi Pons , Marius Miron

Audio fingerprinting techniques have seen great advances in recent years, enabling accurate and fast audio retrieval even in conditions when the queried audio sample has been highly deteriorated or recorded in noisy conditions. Expectedly,…

Information Retrieval · Computer Science 2025-09-26 Kemal Altwlkany , Sead Delalić , Adis Alihodžić , Elmedin Selmanović , Damir Hasić

Recent advancements in audio tokenization have significantly enhanced the integration of audio capabilities into large language models (LLMs). However, audio understanding and generation are often treated as distinct tasks, hindering the…

A noise map facilitates the monitoring of environmental noise pollution in urban areas. However, state-of-the-art techniques for rendering noise maps in urban areas are expensive and rarely updated, as they rely on population and traffic…

Other Computer Science · Computer Science 2013-10-17 Rajib Rana , Chun Tung Chou , Nirupama Bulusu , Salil Kanhere , Wen Hu

The growing sophistication of speech generated by Artificial Intelligence (AI) has introduced new challenges in audio deepfake detection. Text-to-speech (TTS) and voice conversion (VC) technologies can create highly convincing synthetic…

Sound · Computer Science 2026-03-17 Vamshi Nallaguntla , Aishwarya Fursule , Shruti Kshirsagar , Anderson R. Avila

Knowledge of source smartphone corresponding to a document image can be helpful in a variety of applications including copyright infringement, ownership attribution, leak identification and usage restriction. In this letter, we investigate…

Multimedia · Computer Science 2019-06-18 Sharad Joshi , Suraj Saxena , Nitin Khanna

Extraction of the predominant pitch from polyphonic audio is one of the fundamental tasks in the field of music information retrieval and computational musicology. To accomplish this task using machine learning, a large amount of labeled…

Audio and Speech Processing · Electrical Eng. & Systems 2023-04-07 Kavya Ranjan Saxena , Vipul Arora

We introduce PixelPlayer, a system that, by leveraging large amounts of unlabeled videos, learns to locate image regions which produce sounds and separate the input sounds into a set of components that represents the sound from each pixel.…

Computer Vision and Pattern Recognition · Computer Science 2018-10-16 Hang Zhao , Chuang Gan , Andrew Rouditchenko , Carl Vondrick , Josh McDermott , Antonio Torralba

Recognizing speaking in humans is a central task towards understanding social interactions. Ideally, speaking would be detected from individual voice recordings, as done previously for meeting scenarios. However, individual voice recordings…

Computer Vision and Pattern Recognition · Computer Science 2024-03-05 Jose Vargas Quiros , Chirag Raman , Stephanie Tan , Ekin Gedik , Laura Cabrera-Quiros , Hayley Hung
‹ Prev 1 2 3 10 Next ›