Related papers: PIXHELL: When Pixels Learn to Scream
Both acoustic and visual information influence human perception of speech. For this reason, the lack of audio in a video sequence determines an extremely low speech intelligibility for untrained lip readers. In this paper, we present a way…
Spoken language change detection (LCD) refers to detecting language switching points in a multilingual speech signal. Speaker change detection (SCD) refers to locating the speaker change points in a multispeaker speech signal. The objective…
Lipreading has a lot of potential applications such as in the domain of surveillance and video conferencing. Despite this, most of the work in building lipreading systems has been limited to classifying silent videos into classes…
Video generation is one of the most challenging tasks in Machine Learning and Computer Vision fields of study. In this paper, we tackle the text to video generation problem, which is a conditional form of video generation. Humans can…
Cued Speech (CS) is an advanced visual phonetic encoding system that integrates lip reading with hand codings, enabling people with hearing impairments to communicate efficiently. CS video generation aims to produce specific lip and gesture…
Contactless excitation and detection of high harmonic acoustic overtones in a thin insulator single crystal are described using radio frequency spectroscopy techniques. Single crystal [001] silicon wafer samples were investigated, one side…
Lighthill's theory of sound generation was developed to calculate acoustic radiation from a narrow region of turbulent flow embedded in an infinite homogeneous fluid. The theory is extended to include a simple model of non-isothermal medium…
Lip-to-speech involves generating a natural-sounding speech synchronized with a soundless video of a person talking. Despite recent advances, current methods still cannot produce high-quality speech with high levels of intelligibility for…
An ionizing particle passing through a liquid generates acoustic signals via local heat deposition. We delve into modeling such acoustic signals in the case of a single particle that interacts with the liquid electromagnetically in a…
Though visible light communication (VLC) systems are contained to a given room, ensuring their security amongst users in a room is essential. In this paper, the design of artificial noise (AN) to enhance physical layer security in VLC…
Since facial actions such as lip movements contain significant information about speech content, it is not surprising that audio-visual speech enhancement methods are more accurate than their audio-only counterparts. Yet, state-of-the-art…
High resolution digital micro-mirror devices (DMD) make it possible to produce nearly arbitrary light fields with high accuracy, reproducibility and low optical aberrations. However, using these devices to trap and manipulate ultracold…
Understanding the difference between universal low-temperature properties of amorphous and crystalline solids requires an explanation of the stronger damping of long-wavelength phonons in amorphous solids. A longstanding sound attenuation…
Readout noise is a critical parameter for characterizing the performance of charge-coupled devices (CCDs), which can be greatly reduced by the correlated double sampling (CDS) circuit. However, conventional CDS circuit inevitably introduces…
In the present paper, we investigate the properties of the sound generated by rubbing two objects. It is clear that the sound is generated because of the rubbing between the contacting rough surfaces of the objects. A model is presented to…
The advancements in audio generative models have opened up new challenges in their responsible disclosure and the detection of their misuse. In response, we introduce a method to watermark latent generative models by a specific watermarking…
We present a class of sonic meta-screens for manipulating air-borne acoustic waves at ultrasonic or audible frequencies. Our screens consist of periodic arrangements of air bubbles in water or possibly embedded in a soft elastic matrix.…
This paper investigates a novel task of talking face video generation solely from speeches. The speech-to-video generation technique can spark interesting applications in entertainment, customer service, and human-computer-interaction…
We are witnessing a revolution in conditional image synthesis with the recent success of large scale text-to-image generation methods. This success also opens up new opportunities in controlling the generation and editing process using…
High frequency thickness mode ultrasound is an energy-efficient way to atomize high-viscosity fluid at high flow rate into fine aerosol mists of micron-sized droplet distributions. However the complex physics of the atomization process is…