中文
相关论文

相关论文: PIXHELL: When Pixels Learn to Scream

200 篇论文

Traditional RGB-based speech generation faces Temporal Granularity Mismatch since fixed camera exposure times inevitably blur the high-frequency articulatory transients essential for rendering emotional speech. To break this ceiling, we…

多媒体 · 计算机科学 2026-05-27 Jingping Fang , Lin Chen , Chenyang Xu , Tong Zhao , Weidong Cai , Xiaoming Chen

Simulation involves predicting responses of a physical system. In this article, we simulate opto-acoustic signals generated in a three-dimensional volume due to absorption of an optical pulse. A separable computational model is developed…

计算物理 · 物理学 2021-07-22 Jason Zalev , Michael C. Kolios

Sound can move particles. A good example of this phenomenon is the Chladni plate, in which an acoustic wave is induced in a metallic plate and particles migrate to the nodes of the acoustic wave. For several years, acoustophoresis has been…

When a drop impacts a deep pool, it forms a crater which subsequently rebounds. Under certain conditions, a dimple forms at the crater bottom, which pinches off to entrap a small bubble. The oscillation of this entrapped bubble is the…

流体动力学 · 物理学 2025-12-11 Zi Qiang Yang , Yuan Si Tian , Er Qiang Li , Sigurður Tryggvi Thoroddsen

A liquid can be used to represent signals, actuate mechanical computing devices and to modify signals via chemical reactions. We give a brief overview of liquid based computing devices developed over hundreds of years. These include…

新兴技术 · 计算机科学 2018-11-27 Andrew Adamatzky

Humans are able to segment images effortlessly without supervision using perceptual grouping. Here, we propose a counter-intuitive computational approach to solving unsupervised perceptual grouping and segmentation: that they arise because…

计算机视觉与模式识别 · 计算机科学 2024-10-24 Ben Lonnqvist , Zhengqing Wu , Michael H. Herzog

In this paper we present ChirpCast, a system for broadcasting network access keys to laptops ultrasonically. This work explores several modulation techniques for sending and receiving data using sound waves through commodity speakers and…

网络与互联网体系结构 · 计算机科学 2015-08-31 Francis Iannacci , Yanping Huang

While the LED panels used in virtual production systems can display vibrant imagery with a wide color gamut, they produce problematic color shifts when used as lighting due to their peaky spectral output from narrow-band red, green, and…

计算机视觉与模式识别 · 计算机科学 2022-07-05 Chloe LeGendre , Lukas Lepicovsky , Paul Debevec

Spoken language change detection (LCD) refers to identifying the language transitions in a code-switched utterance. Similarly, identifying the speaker transitions in a multispeaker utterance is known as speaker change detection (SCD). Since…

音频与语音处理 · 电气工程与系统科学 2024-06-19 Jagabandhu Mishra , S. R. Mahadeva Prasanna

How does audio describe the world around us? In this work, we propose a method for generating images of visual scenes from diverse in-the-wild sounds. This cross-modal generation task is challenging due to the significant information gap…

计算机视觉与模式识别 · 计算机科学 2024-12-10 Kim Sung-Bin , Arda Senocak , Hyunwoo Ha , Tae-Hyun Oh

Computer generated holography (CGH) has seen a resurgence in recent years due, in part, to the rise of virtual and mixed reality systems. The majority of approaches for CGH are based on a sampled Discrete Fourier Transform (DFT) and ignore…

Ultrasonic acoustic fields have recently been used to generate haptic effects on the human skin as well as to levitate small sub-wavelength size particles. Schlieren imaging and background-oriented schlieren techniques can be used for…

医学物理 · 物理学 2018-10-02 Michele Iodice , William Frier , James Wilcox , Ben Long , Orestis Georgiou

Small vibrations observed in video can unveil information beyond what is visual, such as sound and material properties. It is possible to passively record these vibrations when they are visually perceptible, or actively amplify their visual…

图像与视频处理 · 电气工程与系统科学 2026-01-21 Mingxuan Cai , Dekel Galor , Amit Pal Singh Kohli , Jacob L. Yates , Laura Waller

Previous audio generation mainly focuses on specified sound classes such as speech or music, whose form and content are greatly restricted. In this paper, we go beyond specific audio generation by using natural language description as a…

声音 · 计算机科学 2023-05-04 Guangwei Li , Xuenan Xu , Lingfeng Dai , Mengyue Wu , Kai Yu

Nowadays, CAPTCHAs are computer generated tests that human can pass but current computer systems can not. They have common usage in various web services in order to be able to detect a human from computer programs autonomously. In this way,…

机器学习 · 计算机科学 2019-01-09 Ahmet Faruk Cakmak , Muhammet Balcilar

We introduce SeeingSounds, a lightweight and modular framework for audio-to-image generation that leverages the interplay between audio, language, and vision-without requiring any paired audio-visual data or training on visual generative…

Low-resolution quantized imagery, such as pixel art, is seeing a revival in modern applications ranging from video game graphics to digital design and fabrication, where creativity is often bound by a limited palette of elemental units.…

计算机视觉与模式识别 · 计算机科学 2024-10-10 Alexandre Binninger , Olga Sorkine-Hornung

Surface waves on liquids act as a dynamical phase grating for incident light. In this article, we revisit the classical method of probing such waves (wavelengths of the order of mm) as well as inherent properties of liquids and liquid films…

光学 · 物理学 2009-11-11 Tarun Kr. Barik , Partha Roy Chaudhuri , Anushree Roy , Sayan Kar

Scientific CCDs designed in thick high resistivity silicon (Si) are excellent detectors for astronomy, high energy and nuclear physics, and instrumentation. Many applications can benefit from CCDs ultra low noise readout systems. The…

天体物理仪器与方法 · 物理学 2011-07-06 Gustavo Cancelo , Juan Estrada , Guillermo Fernandez Moroni , Ken Treptow , Ted Zmuda , Tom Diehl

We introduce a novel technique for creative audio resynthesis that operates by reworking the concept of granular synthesis at the latent vector level. Our approach creates a "granular codebook" by encoding a source audio corpus into latent…

声音 · 计算机科学 2025-07-28 Nao Tokui , Tom Baker