中文
相关论文

相关论文: PIXHELL: When Pixels Learn to Scream

200 篇论文

Gestures are essential for enhancing co-speech communication, offering visual emphasis and complementing verbal interactions. While prior work has concentrated on point-level motion or fully supervised data-driven methods, we focus on…

计算机视觉与模式识别 · 计算机科学 2025-03-28 Jiahui Chen , Yang Huan , Runhua Shi , Chanfan Ding , Xiaoqi Mo , Siyu Xiong , Yinong He

Nonlinear optical processes are vital for fields including telecommunications, signal processing, data storage, spectroscopy, sensing, and imaging. As an independent research area, nonlinear optics began with the invention of the laser,…

光学 · 物理学 2018-11-15 Ivan S. Maksymov , Andrew D. Greentree

Visible Light Communication (VLC) using Light Emitting Diodes (LEDs) has gained attention due to its low power consumption, long lifetime, and fast response. However, VLC suffers from optical noise generated by ambient light sources such as…

网络与互联网体系结构 · 计算机科学 2026-02-23 Wataru Uemura , Takumi Hamano

We devise a cascade GAN approach to generate talking face video, which is robust to different face shapes, view angles, facial characteristics, and noisy audio conditions. Instead of learning a direct mapping from audio to video frames, we…

计算机视觉与模式识别 · 计算机科学 2019-05-13 Lele Chen , Ross K. Maddox , Zhiyao Duan , Chenliang Xu

Auditory display is concerned with the use of non-speech sound to communicate information. If the term seems at first oxymoronic, then consider auditory display as an activity of perceptualization, that is, the process of making perceptible…

人机交互 · 计算机科学 2013-11-25 Paul Vickers

Generating accurate sounds for complex audio-visual scenes is challenging, especially in the presence of multiple objects and sound sources. In this paper, we propose an {\em interactive object-aware audio generation} model that grounds…

计算机视觉与模式识别 · 计算机科学 2025-06-05 Tingle Li , Baihe Huang , Xiaobin Zhuang , Dongya Jia , Jiawei Chen , Yuping Wang , Zhuo Chen , Gopala Anumanchipalli , Yuxuan Wang

We introduce an audio texture synthesis algorithm based on scattering moments. A scattering transform is computed by iteratively decomposing a signal with complex wavelet filter banks and computing their amplitude envelop. Scattering…

应用统计 · 统计学 2013-11-05 Joan Bruna , Stéphane Mallat

Efficient face detection is critical to provide natural human-robot interactions. However, computer vision tends to involve a large computational load due to the amount of data (i.e. pixels) that needs to be processed in a short amount of…

音频与语音处理 · 电气工程与系统科学 2024-03-19 William Aris , François Grondin

Spectrograms are 2D representations of sound that look very different from the images found in our visual world. And natural images, when played as spectrograms, make unnatural sounds. In this paper, we show that it is possible to…

计算机视觉与模式识别 · 计算机科学 2025-02-06 Ziyang Chen , Daniel Geng , Andrew Owens

How does textual representation of audio relate to the Large Language Model's (LLMs) learning about the audio world? This research investigates the extent to which LLMs can be prompted to generate audio, despite their primary training in…

Using molecular dynamics simulation, we study acoustic resonance in low-temperature glass by applying a small periodic shear at a boundary wall. Shear wave resonance occurs as the frequency $\omega$ approaches $\omega_\ell= \pi…

软凝聚态物质 · 物理学 2017-08-18 Takeshi Kawasaki , Akira Onuki

Photoacoustic effect refers to the acoustic generation induced by laser irradiation, where nanosecond pulsed laser source is normally used to provide instantaneous heating and thermoelastic expansion of the sample. More generally,…

应用物理 · 物理学 2022-06-08 Yiyun Wang , Qingyuan Shi , Yuting Shen , Yifan Liu , Fei Gao

There has been a growing interest in the task of generating sound for silent videos, primarily because of its practicality in streamlining video post-production. However, existing methods for video-sound generation attempt to directly…

多媒体 · 计算机科学 2024-04-04 Zhifeng Xie , Shengye Yu , Qile He , Mengtian Li

Low power actuation of sessile droplets is of primary interest for portable or hybrid lab-on-a-chip and harmless manipulation of biofluids. In this paper, we show that the acoustic power required to move or deform droplets via surface…

流体动力学 · 物理学 2015-06-04 Michael Baudoin , Philippe Brunet , Olivier Bou Matar , Etienne Herth

This short paper introduces a workflow for generating realistic soundscapes for visual media. In contrast to prior work, which primarily focus on matching sounds for on-screen visuals, our approach extends to suggesting sounds that may not…

声音 · 计算机科学 2023-11-10 David Chuan-En Lin , Nikolas Martelaro

This work presents CLIPDraw, an algorithm that synthesizes novel drawings based on natural language input. CLIPDraw does not require any training; rather a pre-trained CLIP language-image encoder is used as a metric for maximizing…

计算机视觉与模式识别 · 计算机科学 2021-06-29 Kevin Frans , L. B. Soros , Olaf Witkowski

Recently, perceptual image compression has achieved significant advancements, delivering high visual quality at low bitrates for natural images. However, for screen content, existing methods often produce noticeable artifacts when…

计算机视觉与模式识别 · 计算机科学 2025-05-12 Tongda Xu , Jiahao Li , Bin Li , Yan Wang , Ya-Qin Zhang , Yan Lu

Recent advances in visually-induced audio generation are based on sampling short, low-fidelity, and one-class sounds. Moreover, sampling 1 second of audio from the state-of-the-art model takes minutes on a high-end GPU. In this work, we…

计算机视觉与模式识别 · 计算机科学 2021-10-19 Vladimir Iashin , Esa Rahtu

Even though they differ in the physical domain, digital video and audio share many characteristics. Both are temporal data streams often stored in buffers with 8-bit values. This paper investigates a method for creating harmonic sounds with…

人机交互 · 计算机科学 2016-03-01 Carl Thomé

Generating realistic audio effects for movies and other media is a challenging task that is accomplished today primarily through physical techniques known as Foley art. Foley artists create sounds with common objects (e.g., boxing gloves,…

声音 · 计算机科学 2023-08-25 Matthew Martel , Jackson Wagner