English
Related papers

Related papers: PIXHELL: When Pixels Learn to Scream

200 papers

Existing audio analysis methods generally first transform the audio stream to spectrogram, and then feed it into CNN for further analysis. A standard CNN recognizes specific visual patterns over feature map, then pools for high-level…

Sound · Computer Science 2023-03-16 Yulin Pan , Xiangteng He , Biao Gong , Yuxin Peng , Yiliang Lv

The ability of the auditory system to perceive the fundamental frequency of a sound even when this frequency is removed from the stimulus is an interesting phenomenon related to the pitch of complex sounds. This capability is known as…

One way of expressing an environmental sound is using vocal imitations, which involve the process of replicating or mimicking the rhythm and pitch of sounds by voice. We can effectively express the features of environmental sounds, such as…

We present a method that generates expressive talking heads from a single facial image with audio as the only input. In contrast to previous approaches that attempt to learn direct mappings from audio to raw pixels or points for creating…

Computer Vision and Pattern Recognition · Computer Science 2021-02-26 Yang Zhou , Xintong Han , Eli Shechtman , Jose Echevarria , Evangelos Kalogerakis , Dingzeyu Li

In this paper, we present Mixels, programmable magnetic pixels that can be rapidly fabricated using an electromagnetic printhead mounted on an off-the-shelve 3-axis CNC machine. The ability to program magnetic material pixel-wise with…

Human-Computer Interaction · Computer Science 2022-08-09 Martin Nisser , Yashaswini Makaram , Lucian Covarrubias , Amadou Bah , Faraz Faruqi , Ryo Suzuki , Stefanie Mueller

Sound attenuation and internal friction coefficients are calculated for a realistic model of amorphous silicon. It is found that, contrary to previous views, thermal vibrations can induce sound attenuation at ultrasonic and hypersonic…

Condensed Matter · Physics 2009-10-31 Jaroslav Fabian , Philip B. Allen

Talking face generation aims to synthesize a sequence of face images that correspond to a clip of speech. This is a challenging task because face appearance variation and semantics of speech are coupled together in the subtle movements of…

Computer Vision and Pattern Recognition · Computer Science 2019-04-24 Hang Zhou , Yu Liu , Ziwei Liu , Ping Luo , Xiaogang Wang

This paper reports the phenomenon of resonance weakening and streaming onset in two phase acoustofluidics by performing numerical simulations of a capillary droplet suspended in a microfluidic chamber. The simulations show that depending on…

Fluid Dynamics · Physics 2017-03-24 Fabio Garofalo

Pixel detectors typically display pixel-to-pixel gain variation of a few percent which result in reduced spectroscopic performance. We have developed a calibration method which relies on cross-correlating histograms of many pixel pairs and…

Instrumentation and Detectors · Physics 2019-07-17 G. Blaj , G. Haller , C. J. Kenney

This paper presents a novel approach for the automatic generation of Cued Speech (ACSG), a visual communication system used by people with hearing impairment to better elicit the spoken language. We explore transfer learning strategies by…

Computation and Language · Computer Science 2025-01-10 Sanjana Sankar , Martin Lenglet , Gerard Bailly , Denis Beautemps , Thomas Hueber

Denoising diffusion models are widely used for high-quality image and video generation. Their performance depends on noise schedules, which define the distribution of noise levels applied during training and the sequence of noise levels…

Computer Vision and Pattern Recognition · Computer Science 2026-05-12 Carlos Esteves , Ameesh Makadia

A detailed analysis of the use of an optical cavity to enhance picosecond ultrasonic signals is presented. The optical cavity is formed between a distributed Bragg reflector (DBR) and the metal thin film samples to be studied. Experimental…

Materials Science · Physics 2015-05-13 Yanqiu Li , Qian Miao , Arto Nurmikko , Humphrey Maris

Understanding the relationship between the auditory and visual signals is crucial for many different applications ranging from computer-generated imagery (CGI) and video editing automation to assisting people with hearing or visual…

Computer Vision and Pattern Recognition · Computer Science 2020-11-17 Ravindra Yadav , Ashish Sardana , Vinay P Namboodiri , Rajesh M Hegde

The rapid advancement of generative models has made real and synthetic images increasingly indistinguishable. Although extensive efforts have been devoted to detecting AI-generated images, out-of-distribution generalization remains a…

Computer Vision and Pattern Recognition · Computer Science 2026-04-13 Ziqiang Li , Jiazhen Yan , Fan Wang , Kai Zeng , Zhangjie Fu

Latent Consistency Distillation (LCD) has emerged as a promising paradigm for efficient text-to-image synthesis. By distilling a latent consistency model (LCM) from a pre-trained teacher latent diffusion model (LDM), LCD facilitates the…

Computer Vision and Pattern Recognition · Computer Science 2024-10-10 Jiachen Li , Weixi Feng , Wenhu Chen , William Yang Wang

The presence of a corresponding talking face has been shown to significantly improve speech intelligibility in noisy conditions and for hearing impaired population. In this paper, we present a system that can generate landmark points of a…

Computer Vision and Pattern Recognition · Computer Science 2018-04-24 Sefik Emre Eskimez , Ross K Maddox , Chenliang Xu , Zhiyao Duan

The longitudinal oscillations of air columns composed of contractions and rarefaction make up sound. Sound amplification is widely used in medical, electronic and communication fields. A simplistic technique for producing and amplifying can…

Applied Physics · Physics 2025-02-06 Md Hossen Mondal , Ramkrishna A. Joshi

Emulsion droplets trapped in an ultrasonic levitator behave in two ways that solid spheres do not: (1) Individual droplets spin rapidly about an axis parallel to the trapping plane, and (2) coaxially spinning droplets form long chains…

Soft Condensed Matter · Physics 2021-07-02 Mohammed A. Abdelaziz , Jairo A. Diaz , Jean-Luc Aider , David J. Pine , David G. Grier , Mauricio Hoyos

Most existing image compression approaches perform transform coding in the pixel space to reduce its spatial redundancy. However, they encounter difficulties in achieving both high-realism and high-fidelity at low bitrate, as the…

Computer Vision and Pattern Recognition · Computer Science 2025-12-24 Zhaoyang Jia , Jiahao Li , Bin Li , Houqiang Li , Yan Lu

We propose a method leveraging the naturally time-related expressivity of our voice to control an animation composed of a set of short events. The user records itself mimicking onomatopoeia sounds such as "Tick", "Pop", or "Chhh" which are…

Graphics · Computer Science 2019-10-21 Adrien Nivaggioli , Damien Rohmer