Related papers: PIXHELL: When Pixels Learn to Scream
This paper presents an audio-visual approach for voice separation which produces state-of-the-art results at a low latency in two scenarios: speech and singing voice. The model is based on a two-stage network. Motion cues are obtained with…
Complex scattering potentials can admit scattering states that behave exactly like a zero-width resonance. Their energy is what mathematicians call a spectral singularity. This phenomenon admits optical realizations in the form of lasing at…
A simple and robust experiment demonstrating computational ghost imaging with structured illumination and a single-pixel detector has been performed. Our experimental setup utilizes a general computer for generating pseudo-randomly patterns…
Acoustic lenses are employed in a variety of applications, from biomedical imaging and surgery, to defense systems, but their performance is limited by their linear operational envelope and complexity. Here we show a dramatic focusing…
Unlike existing methods that rely on source images as appearance references and use source speech to generate motion, this work proposes a novel approach that directly extracts information from the speech, addressing key challenges in…
A nonlinear optical medium results by the collective orientation of liquid crystal molecules tightly coupled to a transparent photoconductive layer. We show that such a medium can give a large gain, thus, if inserted in a ring cavity, it…
Vocoders received renewed attention as main components in statistical parametric text-to-speech (TTS) synthesis and speech transformation systems. Even though there are vocoding techniques give almost accepted synthesized speech, their high…
Audio-driven talking head generation has drawn growing attention. To produce talking head videos with desired facial expressions, previous methods rely on extra reference videos to provide expression information, which may be difficult to…
Large-scale pre-trained image-text models demonstrate remarkable versatility across diverse tasks, benefiting from their robust representational capabilities and effective multimodal alignment. We extend the application of these models,…
Computer-generated holographic (CGH) displays show great potential and are emerging as the next-generation displays for augmented and virtual reality, and automotive heads-up displays. One of the critical problems harming the wide adoption…
Recently, the application of diffusion probabilistic models has advanced speech enhancement through generative approaches. However, existing diffusion-based methods have focused on the generation process in high-dimensional waveform or…
When video is shot in noisy environment, the voice of a speaker seen in the video can be enhanced using the visible mouth movements, reducing background noise. While most existing methods use audio-only inputs, improved performance is…
A radically new CCD development by Marconi Applied Technologies has enabled substantial internal gain within the CCD before the signal reaches the output amplifier. With reasonably high gain, sub-electron readout noise levels are achieved…
Ultrafast electron diffraction and phonon-diffuse scattering (UED(S)) experiments make use of photo-induced changes to electron scattering intensity across 2D detectors to report on a very wide range of dynamic structural phenomena in…
In this work, we introduce the novel technique of in-chip drop on demand, which consists in dispensing picoliter to nanoliter drops on demand directly in the liquid-filled channels of a polymer microfluidic chip, at frequencies up to 2.5…
Representing wild sounds as images is an important but challenging task due to the lack of paired datasets between sound and images and the significant differences in the characteristics of these two modalities. Previous studies have…
Integrated optical devices may replace bulk crystal or fiber based assemblies with a more compact and controllable photon pair and heralded single photon source and generate quantum light at telecommunications wavelengths. Here, we propose…
Deep learning, with its robust aotomatic feature extraction capabilities, has demonstrated significant success in audio signal processing. Typically, these methods rely on static, pre-collected large-scale datasets for training, performing…
Interactive acoustic auralization allows users to explore virtual acoustic environments in real-time, enabling the acoustic recreation of concert hall or Historical Worship Spaces (HWS) that are either no longer accessible, acoustically…
We use the formerly derived explicit analytical expressions for the conductivity of nanostructured superconductors supercooled below the critical temperature in electric field. Computer simulations reveal that the negative differential…