Related papers: PIXHELL: When Pixels Learn to Scream
Existing audio analysis methods generally first transform the audio stream to spectrogram, and then feed it into CNN for further analysis. A standard CNN recognizes specific visual patterns over feature map, then pools for high-level…
The ability of the auditory system to perceive the fundamental frequency of a sound even when this frequency is removed from the stimulus is an interesting phenomenon related to the pitch of complex sounds. This capability is known as…
One way of expressing an environmental sound is using vocal imitations, which involve the process of replicating or mimicking the rhythm and pitch of sounds by voice. We can effectively express the features of environmental sounds, such as…
We present a method that generates expressive talking heads from a single facial image with audio as the only input. In contrast to previous approaches that attempt to learn direct mappings from audio to raw pixels or points for creating…
In this paper, we present Mixels, programmable magnetic pixels that can be rapidly fabricated using an electromagnetic printhead mounted on an off-the-shelve 3-axis CNC machine. The ability to program magnetic material pixel-wise with…
Sound attenuation and internal friction coefficients are calculated for a realistic model of amorphous silicon. It is found that, contrary to previous views, thermal vibrations can induce sound attenuation at ultrasonic and hypersonic…
Talking face generation aims to synthesize a sequence of face images that correspond to a clip of speech. This is a challenging task because face appearance variation and semantics of speech are coupled together in the subtle movements of…
This paper reports the phenomenon of resonance weakening and streaming onset in two phase acoustofluidics by performing numerical simulations of a capillary droplet suspended in a microfluidic chamber. The simulations show that depending on…
Pixel detectors typically display pixel-to-pixel gain variation of a few percent which result in reduced spectroscopic performance. We have developed a calibration method which relies on cross-correlating histograms of many pixel pairs and…
This paper presents a novel approach for the automatic generation of Cued Speech (ACSG), a visual communication system used by people with hearing impairment to better elicit the spoken language. We explore transfer learning strategies by…
Denoising diffusion models are widely used for high-quality image and video generation. Their performance depends on noise schedules, which define the distribution of noise levels applied during training and the sequence of noise levels…
A detailed analysis of the use of an optical cavity to enhance picosecond ultrasonic signals is presented. The optical cavity is formed between a distributed Bragg reflector (DBR) and the metal thin film samples to be studied. Experimental…
Understanding the relationship between the auditory and visual signals is crucial for many different applications ranging from computer-generated imagery (CGI) and video editing automation to assisting people with hearing or visual…
The rapid advancement of generative models has made real and synthetic images increasingly indistinguishable. Although extensive efforts have been devoted to detecting AI-generated images, out-of-distribution generalization remains a…
Latent Consistency Distillation (LCD) has emerged as a promising paradigm for efficient text-to-image synthesis. By distilling a latent consistency model (LCM) from a pre-trained teacher latent diffusion model (LDM), LCD facilitates the…
The presence of a corresponding talking face has been shown to significantly improve speech intelligibility in noisy conditions and for hearing impaired population. In this paper, we present a system that can generate landmark points of a…
The longitudinal oscillations of air columns composed of contractions and rarefaction make up sound. Sound amplification is widely used in medical, electronic and communication fields. A simplistic technique for producing and amplifying can…
Emulsion droplets trapped in an ultrasonic levitator behave in two ways that solid spheres do not: (1) Individual droplets spin rapidly about an axis parallel to the trapping plane, and (2) coaxially spinning droplets form long chains…
Most existing image compression approaches perform transform coding in the pixel space to reduce its spatial redundancy. However, they encounter difficulties in achieving both high-realism and high-fidelity at low bitrate, as the…
We propose a method leveraging the naturally time-related expressivity of our voice to control an animation composed of a set of short events. The user records itself mimicking onomatopoeia sounds such as "Tick", "Pop", or "Chhh" which are…