中文
相关论文

相关论文: PIXHELL: When Pixels Learn to Scream

200 篇论文

At low temperatures, elementary excitations of a one-dimensional quantum liquid form a gas that can move as a whole with respect to the center of mass of the system. This internal motion attenuates at exponentially long time scales. As a…

介观与纳米尺度物理 · 物理学 2018-11-07 K. A. Matveev , A. V. Andreev

The anomalous low-temperature properties of glasses arise from intrinsic excitable entities, so-called tunneling Two-Level-Systems (TLS), whose microscopic nature has been baffling solid-state physicists for decades. TLS have become…

We have constructed and characterised an instrument to study gravitationally bouncing droplets of fluid, subjected to periodic driving force. Our system incorporates a droplet printer that enables an on-demand computer controlled deposition…

流体动力学 · 物理学 2024-03-12 Tapio Simula

Owing to the power of vision-language foundation models, e.g., CLIP, the area of image synthesis has seen recent important advances. Particularly, for style transfer, CLIP enables transferring more general and abstract styles without…

计算机视觉与模式识别 · 计算机科学 2023-11-07 Zipeng Xu , Songlong Xing , Enver Sangineto , Nicu Sebe

Speechreading is a notoriously difficult task for humans to perform. In this paper we present an end-to-end model based on a convolutional neural network (CNN) for generating an intelligible acoustic speech signal from silent video frames…

计算机视觉与模式识别 · 计算机科学 2017-01-10 Ariel Ephrat , Shmuel Peleg

Vocoders are models capable of transforming a low-dimensional spectral representation of an audio signal, typically the mel spectrogram, to a waveform. Modern speech generation pipelines use a vocoder as their final component. Recent…

声音 · 计算机科学 2022-08-29 Bruno Di Giorgi , Mark Levy , Richard Sharp

The so-called Locally Resonant Acoustic Metamaterials (LRAM) are considered for the design of specifically engineered devices capable of stopping waves from propagating in certain frequency regions (bandgaps), this making them applicable…

计算工程、金融与科学 · 计算机科学 2021-08-16 D. Roca , D. Yago , J. Cante , O. Lloberas-Valls , J. Oliver

Sound waves cause small vibrations in nearby objects. A few techniques exist in the literature that can extract sound from video. In this paper we study local vibration patterns at different image locations. We show that different locations…

计算机视觉与模式识别 · 计算机科学 2019-07-16 Mohammad Amin Shabani , Laleh Samadfam , Mohammad Amin Sadeghi

We present SoundLoCD, a novel text-to-sound generation framework, which incorporates a LoRA-based conditional discrete contrastive latent diffusion model. Unlike recent large-scale sound generation models, our model can be efficiently…

声音 · 计算机科学 2024-05-27 Xinlei Niu , Jing Zhang , Christian Walder , Charles Patrick Martin

This paper introduces a method for using LED-based environmental lighting to produce visually imperceptible watermarks for consumer cameras. Our approach optimizes an LED light source's spectral profile to be minimally visible to the human…

计算机视觉与模式识别 · 计算机科学 2025-11-18 Hodaka Kawachi , Tomoya Nakamura , Hiroaki Santo , SaiKiran Kumar Tedla , Trevor Dalton Canham , Yasushi Yagi , Michael S. Brown

We present an end-to-end binaural audio rendering approach (Listen2Scene) for virtual reality (VR) and augmented reality (AR) applications. We propose a novel neural-network-based binaural sound propagation method to generate acoustic…

音频与语音处理 · 电气工程与系统科学 2024-02-09 Anton Ratnarajah , Dinesh Manocha

Acoustical mismatch among training and testing phases degrades outstandingly speech recognition results. This problem has limited the development of real-world nonspecific applications, as testing conditions are highly variant or even…

声音 · 计算机科学 2013-05-13 Rashmi Makhijani , Urmila Shrawankar , V M Thakare

Noise synthesis is a challenging low-level vision task aiming to generate realistic noise given a clean image along with the camera settings. To this end, we propose an effective generative model which utilizes clean features as guidance…

计算机视觉与模式识别 · 计算机科学 2024-06-14 Mingyang Song , Yang Zhang , Tunç O. Aydın , Elham Amin Mansour , Christopher Schroers

We present a method for automatically producing human-like vocal imitations of sounds: the equivalent of "sketching," but for auditory rather than visual representation. Starting with a simulated model of the human vocal tract, we first try…

图形学 · 计算机科学 2024-09-23 Matthew Caren , Kartik Chandra , Joshua B. Tenenbaum , Jonathan Ragan-Kelley , Karima Ma

In recent works, a flow-based neural vocoder has shown significant improvement in real-time speech generation task. The sequence of invertible flow operations allows the model to convert samples from simple distribution to audio samples.…

音频与语音处理 · 电气工程与系统科学 2020-08-18 Hyun-Wook Yoon , Sang-Hoon Lee , Hyeong-Rae Noh , Seong-Whan Lee

We demonstrate how conditional generation from diffusion models can be used to tackle a variety of realistic tasks in the production of music in 44.1kHz stereo audio with sampling-time guidance. The scenarios we consider include…

声音 · 计算机科学 2023-12-06 Mark Levy , Bruno Di Giorgi , Floris Weers , Angelos Katharopoulos , Tom Nickson

Readout chips of hybrid pixel detectors use a low power amplifier and threshold discrimination to process charge deposited in semiconductor sensors. Due to transistor mismatch each pixel circuit needs to be calibrated individually to…

仪器与探测器 · 物理学 2017-08-23 Timon Heim , Maurice Garcia-Sciveres

Contrastive models like CLIP have been shown to learn robust representations of images that capture both semantics and style. To leverage these representations for image generation, we propose a two-stage model: a prior that generates a…

计算机视觉与模式识别 · 计算机科学 2022-04-14 Aditya Ramesh , Prafulla Dhariwal , Alex Nichol , Casey Chu , Mark Chen

Natural language is widely used to describe, prompt, and control audio systems, but rarely serves as the representation carrying audio itself. We introduce lexical acoustic coding (LAC), a framework in which pre-trained LLM sender and…

机器学习 · 计算机科学 2026-05-12 Emanuele Rossi , Emanuele Rodolà

This paper describes and validates for the first time the dynamic modelling of Liquid Crystal (LC)-based planar multi-resonant cells, as well as its use as bias signals synthesis tool to improve their reconfigurability time. The dynamic LC…