中文
相关论文

相关论文: PIXHELL: When Pixels Learn to Scream

200 篇论文

In this work, we report on inelastic X-ray scattering experiments combined with the molecular dynamics simulations on deeply supercritical Ar. The presented results unveil the mechanism and regimes of sound propagation in the liquid matter…

Tactile displays that lend tangible form to digital content could transform computing interactions. However, achieving the resolution, speed, and dynamic range needed for perceptual fidelity remains challenging. We present a tactile display…

新兴技术 · 计算机科学 2026-01-27 Max Linnander , Dustin Goetz , Gregory Reardon , Vijay Kumar , Elliot Hawkes , Yon Visell

Pixel diffusion generates images directly in pixel space, avoiding the VAE artifacts and representational bottlenecks of two-stage latent diffusion. Recent JiT further simplifies pixel diffusion with x-prediction, where the model predicts…

计算机视觉与模式识别 · 计算机科学 2026-05-08 Zehong Ma , Ruihan Xu , Shiliang Zhang

Recent work showed the possibility of building open-vocabulary large language models (LLMs) that directly operate on pixel representations. These models are implemented as autoencoders that reconstruct masked patches of rendered text.…

计算与语言 · 计算机科学 2024-02-27 Yintao Tai , Xiyang Liao , Alessandro Suglia , Antonio Vergari

We use an optical centrifuge to deposit a controllable amount of rotational energy into dense molecular ensembles. Subsequent rotation-translation energy transfer, mediated by thermal collisions, results in the localized heating of the gas…

化学物理 · 物理学 2015-06-23 A. A. Milner , A. Korobenko , V. Milner

In recent years, image generation has shown a great leap in performance, where diffusion models play a central role. Although generating high-quality images, such models are mainly conditioned on textual descriptions. This begs the…

声音 · 计算机科学 2023-05-23 Guy Yariv , Itai Gat , Lior Wolf , Yossi Adi , Idan Schwartz

We propose "Insect-Computer Hybrid Speaker", which enables us to make musics made from combinations of computer and insects. Lots of studies have proposed methods and interfaces for controlling insects and obtaining feedback. However, there…

人机交互 · 计算机科学 2025-04-24 Yuga Tsukuda , Naoto Nishida , Jun Lu , Yoichi Ochiai

Most vision-language systems are static observers: they describe pixels, do not act, and cannot safely improve under shift. This passivity limits generalizable, physically grounded visual intelligence. Learning through action, not static…

计算机视觉与模式识别 · 计算机科学 2026-03-27 Yunpeng Zhou

We present a user-friendly and versatile experimental technique that generates sub-kilohertz sinusoidal oscillatory flows within microchannels. The method involves the direct interfacing of microfluidic tubing with a loudspeaker diaphragm…

流体动力学 · 物理学 2020-08-03 Giridar Vishwanathan , Gabriel Juarez

The introduction of audio latent diffusion models possessing the ability to generate realistic sound clips on demand from a text description has the potential to revolutionize how we work with audio. In this work, we make an initial attempt…

音频与语音处理 · 电气工程与系统科学 2023-10-17 Dimitrios Bralios , Gordon Wichern , François G. Germain , Zexu Pan , Sameer Khurana , Chiori Hori , Jonathan Le Roux

Image-text contrastive models like CLIP have wide applications in zero-shot classification, image-text retrieval, and transfer learning. However, they often struggle on compositional visio-linguistic tasks (e.g., attribute-binding or…

计算机视觉与模式识别 · 计算机科学 2024-07-02 Samyadeep Basu , Shell Xu Hu , Maziar Sanjabi , Daniela Massiceti , Soheil Feizi

We present a simple nearest-neighbor (NN) approach that synthesizes high-frequency photorealistic images from an "incomplete" signal such as a low-resolution image, a surface normal map, or edges. Current state-of-the-art deep generative…

计算机视觉与模式识别 · 计算机科学 2017-08-18 Aayush Bansal , Yaser Sheikh , Deva Ramanan

This paper proposes an expressive singing voice synthesis system by introducing explicit vibrato modeling and latent energy representation. Vibrato is essential to the naturalness of synthesized sound, due to the inherent characteristics of…

声音 · 计算机科学 2022-11-03 Yingjie Song , Wei Song , Wei Zhang , Zhengchen Zhang , Dan Zeng , Zhi Liu , Yang Yu

Generative models are successfully used for image synthesis in the recent years. But when it comes to other modalities like audio, text etc little progress has been made. Recent works focus on generating audio from a generative model in an…

计算机视觉与模式识别 · 计算机科学 2018-09-30 Chae Young Lee , Anoop Toffy , Gue Jun Jung , Woo-Jin Han

In various Computer Vision and Signal Processing applications, noise is typically perceived as a drawback of the image capturing system that ought to be removed. We, on the other hand, claim that image noise, just as texture, is important…

计算机视觉与模式识别 · 计算机科学 2018-03-28 Renata Khasanova , Jan Wassenberg , Jyrki Alakuijala

Departing from traditional communication theory where decoding algorithms are assumed to perform without error, a system where noise perturbs both computational devices and communication channels is considered here. This paper studies…

信息论 · 计算机科学 2010-05-31 Lav R. Varshney

Objects make distinctive sounds when they are hit or scratched. These sounds reveal aspects of an object's material properties, as well as the actions that produced them. In this paper, we propose the task of predicting what sound an object…

计算机视觉与模式识别 · 计算机科学 2016-05-03 Andrew Owens , Phillip Isola , Josh McDermott , Antonio Torralba , Edward H. Adelson , William T. Freeman

Error-control-coding (ECC) techniques are widely used in modern digital communication systems to minimize the effect of noisy channels on the quality of received signals. Motivated by the fact that both communication and imaging can be…

图像与视频处理 · 电气工程与系统科学 2018-09-20 Xiaopeng Wang , Zunwang Bo , Zihuai Lin , Wenlin Gong , Branka Vucetic , Shensheng Han

Procedural noise is a fundamental component of computer graphics pipelines, offering a flexible way to generate textures that exhibit "natural" random variation. Many different types of noise exist, each produced by a separate algorithm. In…

Learning a new language involves constantly comparing speech productions with reference productions from the environment. Early in speech acquisition, children make articulatory adjustments to match their caregivers' speech. Grownup…

音频与语音处理 · 电气工程与系统科学 2022-07-01 Talia Ben-Simon , Felix Kreuk , Faten Awwad , Jacob T. Cohen , Joseph Keshet