English
Related papers

Related papers: PIXHELL: When Pixels Learn to Scream

200 papers

At low temperatures, elementary excitations of a one-dimensional quantum liquid form a gas that can move as a whole with respect to the center of mass of the system. This internal motion attenuates at exponentially long time scales. As a…

Mesoscale and Nanoscale Physics · Physics 2018-11-07 K. A. Matveev , A. V. Andreev

The anomalous low-temperature properties of glasses arise from intrinsic excitable entities, so-called tunneling Two-Level-Systems (TLS), whose microscopic nature has been baffling solid-state physicists for decades. TLS have become…

We have constructed and characterised an instrument to study gravitationally bouncing droplets of fluid, subjected to periodic driving force. Our system incorporates a droplet printer that enables an on-demand computer controlled deposition…

Fluid Dynamics · Physics 2024-03-12 Tapio Simula

Owing to the power of vision-language foundation models, e.g., CLIP, the area of image synthesis has seen recent important advances. Particularly, for style transfer, CLIP enables transferring more general and abstract styles without…

Computer Vision and Pattern Recognition · Computer Science 2023-11-07 Zipeng Xu , Songlong Xing , Enver Sangineto , Nicu Sebe

Speechreading is a notoriously difficult task for humans to perform. In this paper we present an end-to-end model based on a convolutional neural network (CNN) for generating an intelligible acoustic speech signal from silent video frames…

Computer Vision and Pattern Recognition · Computer Science 2017-01-10 Ariel Ephrat , Shmuel Peleg

Vocoders are models capable of transforming a low-dimensional spectral representation of an audio signal, typically the mel spectrogram, to a waveform. Modern speech generation pipelines use a vocoder as their final component. Recent…

Sound · Computer Science 2022-08-29 Bruno Di Giorgi , Mark Levy , Richard Sharp

The so-called Locally Resonant Acoustic Metamaterials (LRAM) are considered for the design of specifically engineered devices capable of stopping waves from propagating in certain frequency regions (bandgaps), this making them applicable…

Computational Engineering, Finance, and Science · Computer Science 2021-08-16 D. Roca , D. Yago , J. Cante , O. Lloberas-Valls , J. Oliver

Sound waves cause small vibrations in nearby objects. A few techniques exist in the literature that can extract sound from video. In this paper we study local vibration patterns at different image locations. We show that different locations…

Computer Vision and Pattern Recognition · Computer Science 2019-07-16 Mohammad Amin Shabani , Laleh Samadfam , Mohammad Amin Sadeghi

We present SoundLoCD, a novel text-to-sound generation framework, which incorporates a LoRA-based conditional discrete contrastive latent diffusion model. Unlike recent large-scale sound generation models, our model can be efficiently…

Sound · Computer Science 2024-05-27 Xinlei Niu , Jing Zhang , Christian Walder , Charles Patrick Martin

This paper introduces a method for using LED-based environmental lighting to produce visually imperceptible watermarks for consumer cameras. Our approach optimizes an LED light source's spectral profile to be minimally visible to the human…

Computer Vision and Pattern Recognition · Computer Science 2025-11-18 Hodaka Kawachi , Tomoya Nakamura , Hiroaki Santo , SaiKiran Kumar Tedla , Trevor Dalton Canham , Yasushi Yagi , Michael S. Brown

We present an end-to-end binaural audio rendering approach (Listen2Scene) for virtual reality (VR) and augmented reality (AR) applications. We propose a novel neural-network-based binaural sound propagation method to generate acoustic…

Audio and Speech Processing · Electrical Eng. & Systems 2024-02-09 Anton Ratnarajah , Dinesh Manocha

Acoustical mismatch among training and testing phases degrades outstandingly speech recognition results. This problem has limited the development of real-world nonspecific applications, as testing conditions are highly variant or even…

Sound · Computer Science 2013-05-13 Rashmi Makhijani , Urmila Shrawankar , V M Thakare

Noise synthesis is a challenging low-level vision task aiming to generate realistic noise given a clean image along with the camera settings. To this end, we propose an effective generative model which utilizes clean features as guidance…

Computer Vision and Pattern Recognition · Computer Science 2024-06-14 Mingyang Song , Yang Zhang , Tunç O. Aydın , Elham Amin Mansour , Christopher Schroers

We present a method for automatically producing human-like vocal imitations of sounds: the equivalent of "sketching," but for auditory rather than visual representation. Starting with a simulated model of the human vocal tract, we first try…

Graphics · Computer Science 2024-09-23 Matthew Caren , Kartik Chandra , Joshua B. Tenenbaum , Jonathan Ragan-Kelley , Karima Ma

In recent works, a flow-based neural vocoder has shown significant improvement in real-time speech generation task. The sequence of invertible flow operations allows the model to convert samples from simple distribution to audio samples.…

Audio and Speech Processing · Electrical Eng. & Systems 2020-08-18 Hyun-Wook Yoon , Sang-Hoon Lee , Hyeong-Rae Noh , Seong-Whan Lee

We demonstrate how conditional generation from diffusion models can be used to tackle a variety of realistic tasks in the production of music in 44.1kHz stereo audio with sampling-time guidance. The scenarios we consider include…

Sound · Computer Science 2023-12-06 Mark Levy , Bruno Di Giorgi , Floris Weers , Angelos Katharopoulos , Tom Nickson

Readout chips of hybrid pixel detectors use a low power amplifier and threshold discrimination to process charge deposited in semiconductor sensors. Due to transistor mismatch each pixel circuit needs to be calibrated individually to…

Instrumentation and Detectors · Physics 2017-08-23 Timon Heim , Maurice Garcia-Sciveres

Contrastive models like CLIP have been shown to learn robust representations of images that capture both semantics and style. To leverage these representations for image generation, we propose a two-stage model: a prior that generates a…

Computer Vision and Pattern Recognition · Computer Science 2022-04-14 Aditya Ramesh , Prafulla Dhariwal , Alex Nichol , Casey Chu , Mark Chen

Natural language is widely used to describe, prompt, and control audio systems, but rarely serves as the representation carrying audio itself. We introduce lexical acoustic coding (LAC), a framework in which pre-trained LLM sender and…

Machine Learning · Computer Science 2026-05-12 Emanuele Rossi , Emanuele Rodolà

This paper describes and validates for the first time the dynamic modelling of Liquid Crystal (LC)-based planar multi-resonant cells, as well as its use as bias signals synthesis tool to improve their reconfigurability time. The dynamic LC…