English
Related papers

Related papers: An audio-only method for advertisement detection i…

200 papers

We tackle the problem of audiovisual scene analysis for weakly-labeled data. To this end, we build upon our previous audiovisual representation learning framework to perform object classification in noisy acoustic environments and integrate…

Computer Vision and Pattern Recognition · Computer Science 2018-11-12 Sanjeel Parekh , Alexey Ozerov , Slim Essid , Ngoc Duong , Patrick Pérez , Gaël Richard

Perceptual ad-blocking is a novel approach that detects online advertisements based on their visual content. Compared to traditional filter lists, the use of perceptual signals is believed to be less prone to an arms race with web…

Cryptography and Security · Computer Science 2019-08-27 Florian Tramèr , Pascal Dupré , Gili Rusak , Giancarlo Pellegrino , Dan Boneh

Large audio-language models (LALMs) exhibit strong zero-shot capabilities in multiple downstream tasks, such as audio question answering (AQA) and abstract reasoning; however, these models still lag behind specialized models for certain…

Sound · Computer Science 2026-03-24 Videet Mehta , Liming Wang , Hilde Kuehne , Rogerio Feris , James R. Glass , M. Jehanzeb Mirza

Generally audio news broadcast on radio is com- posed of music, commercials, news from correspondents and recorded statements in addition to the actual news read by the newsreader. When news transcripts are available, automatic segmentation…

Sound · Computer Science 2014-03-28 Sapna Soni , Ahmed Imran , Sunil Kumar Kopparapu

In this paper, we focus on the task of one-shot sign spotting, i.e. given an example of an isolated sign (query), we want to identify whether/where this sign appears in a continuous, co-articulated sign language video (target). To achieve…

Computer Vision and Pattern Recognition · Computer Science 2021-11-23 Tao Jiang , Necati Cihan Camgoz , Richard Bowden

This article develops a general detection theory for speech analysis based on time-varying autoregressive models, which themselves generalize the classical linear predictive speech analysis framework. This theory leads to a computationally…

Applications · Statistics 2011-08-25 Daniel Rudoy , Thomas F. Quatieri , Patrick J. Wolfe

The analysis, processing, and extraction of meaningful information from sounds all around us is the subject of the broader area of audio analytics. Audio captioning is a recent addition to the domain of audio analytics, a cross-modal…

Audio and Speech Processing · Electrical Eng. & Systems 2023-05-04 Sandeep Kothinti , Dimitra Emmanouilidou

Energy detection is a reliable non-coherent signal processing technology of spectrum sensing of cognitive radio networks, which thanks to its low complexity, no requirement of priori received information and fast sensing ability etc. Since…

Information Theory · Computer Science 2018-06-14 He Huang

We demonstrate the existence of universal adversarial perturbations, which can fool a family of audio classification architectures, for both targeted and untargeted attack scenarios. We propose two methods for finding such perturbations.…

Machine Learning · Computer Science 2020-11-18 Sajjad Abdoli , Luiz G. Hafemann , Jerome Rony , Ismail Ben Ayed , Patrick Cardinal , Alessandro L. Koerich

Children comprise a significant proportion of TV viewers and it is worthwhile to customize the experience for them. However, identifying who is a child in the audience can be a challenging task. Identifying gender and age from audio…

Computation and Language · Computer Science 2018-03-05 Denys Katerenchuk

Active speaker detection is a challenging task in audio-visual scenario understanding, which aims to detect who is speaking in one or more speakers scenarios. This task has received extensive attention as it is crucial in applications such…

Computer Vision and Pattern Recognition · Computer Science 2023-03-09 Junhua Liao , Haihan Duan , Kanghui Feng , Wanbing Zhao , Yanbing Yang , Liangyin Chen

Sound event detection systems are widely used in various applications such as surveillance and environmental monitoring where data is automatically collected, processed, and sent to a cloud for sound recognition. However, this process may…

Sound · Computer Science 2024-01-04 Shayan Gharib , Minh Tran , Diep Luong , Konstantinos Drossos , Tuomas Virtanen

Short-range audio channels have a few distinguishing characteristics: ease of use, low deployment costs, and easy to tune frequencies, to cite a few. Moreover, thanks to their seamless adaptability to the security context, many techniques…

Cryptography and Security · Computer Science 2022-08-09 Maurantonio Caprolu , Savio Sciancalepore , Roberto Di Pietro

Keyword spotting is an important research field because it plays a key role in device wake-up and user interaction on smart devices. However, it is challenging to minimize errors while operating efficiently in devices with limited resources…

Sound · Computer Science 2023-07-06 Byeonggeun Kim , Simyung Chang , Jinkyu Lee , Dooyong Sung

This paper introduces a computational framework designed to delineate gender distribution biases in topics covered by French TV and radio news. We transcribe a dataset of 11.7k hours, broadcasted in 2023 on 21 French channels. A Large…

Computation and Language · Computer Science 2024-07-22 Valentin Pelloin , Lena Dodson , Émile Chapuis , Nicolas Hervé , David Doukhan

The article addresses the problem of detecting presence and location of a small low emission source inside of an object, when the background noise dominates. This problem arises, for instance, in some homeland security applications. The…

Statistics Theory · Mathematics 2015-05-28 Xiaolei Xun , Bani Mallick , Raymond J. Carroll , Peter Kuchment

Online abusive content detection, particularly in low-resource settings and within the audio modality, remains underexplored. We investigate the potential of pre-trained audio representations for detecting abusive language in low-resource…

Computation and Language · Computer Science 2024-12-16 Aditya Narayan Sankaran , Reza Farahbakhsh , Noel Crespi

We present a model for separating a set of voices out of a sound mixture containing an unknown number of sources. Our Attentional Gating Network (AGN) uses a variable attentional context to specify which speakers in the mixture are of…

Sound · Computer Science 2019-05-28 Shariq Mobin , Bruno Olshausen

In this paper, we consider a simple coding scheme for spatial modulation (SM), where the same set of active transmit antennas is repeatedly used over consecutive multiple transmissions. Based on a Gaussian approximation, an approximate…

Information Theory · Computer Science 2019-01-01 Jinho Choi

In this paper we present a non-invasive ambient intelligence framework for the semi-automatic analysis of non-verbal communication applied to the restorative justice field. In particular, we propose the use of computer vision and social…

Human-Computer Interaction · Computer Science 2016-02-22 Víctor Ponce-López , Sergio Escalera , Marc Pérez , Oriol Janés , Xavier Baró