English
Related papers

Related papers: PIXHELL: When Pixels Learn to Scream

200 papers

CLIP is a discriminative model trained to align images and text in a shared embedding space. Due to its multimodal structure, it serves as the backbone of many generative pipelines, where a decoder is trained to map from the shared space…

Computer Vision and Pattern Recognition · Computer Science 2025-06-05 Antonio D'Orazio , Maria Rosaria Briglia , Donato Crisostomi , Dario Loi , Emanuele Rodolà , Iacopo Masi

It has been recently observed that synthetic materials subjected to an external elastic stress give rise to scaling phenomena in the acoustic emission signal. Motivated by this experimental finding we develop a mesoscopic model in order to…

Materials Science · Physics 2008-02-03 Stefano Zapperi , Alessandro Vespignani , H. Eugene Stanley

Existing video generation models excel at producing photo-realistic videos from text or images, but often lack physical plausibility and 3D controllability. To overcome these limitations, we introduce PhysCtrl, a novel framework for…

Computer Vision and Pattern Recognition · Computer Science 2025-11-11 Chen Wang , Chuhao Chen , Yiming Huang , Zhiyang Dou , Yuan Liu , Jiatao Gu , Lingjie Liu

This paper proposes a new research direction for the large family of instrumental musical interfaces where sound is generated using digital granular synthesis, and where interaction and control involve the (fine) operation of stiff, flat…

Human-Computer Interaction · Computer Science 2021-07-28 Staas de Jong

CVD diamond is an attractive material option for LHC vertex detectors because of its strong radiation-hardness causal to its large band gap and strong lattice. In particular, pixel detectors operating close to the interaction point profit…

Instrumentation and Detectors · Physics 2015-06-05 Jieh-Wen Tsung , Miroslav Havranek , Fabian Hügging , Harris Kagan , Hans Krüger , Norbert Wermes

Data-driven approaches hold promise for audio captioning. However, the development of audio captioning methods can be biased due to the limited availability and quality of text-audio data. This paper proposes a SynthAC framework, which…

Sound · Computer Science 2023-09-19 Feiyang Xiao , Qiaoxi Zhu , Jian Guan , Xubo Liu , Haohe Liu , Kejia Zhang , Wenwu Wang

Sonic crystal acoustic screens have been in progressive research and development in the last two decades as a technical solution for mitigating traffic noise. Their behaviour is quite different from that observed in classical barriers, with…

This paper introduces a system that learns to sing new tunes by listening to examples. It extracts sequencing rules from input music and uses these rules to generate new tunes, which are sung by a vocal synthesiser. We developed a method to…

Quantum Physics · Physics 2022-08-09 Eduardo Reck Miranda , Brian N. Siegelwax

We introduce multilayer structures based on phase-change materials for reconfigurable structural color generation. These structures can produce multiple distinct colors within a single pixel. Specifically, we design structures that generate…

We suggest a rasterization pipeline tailored towards the need of head-mounted displays (HMD), where latency and field-of-view requirements pose new challenges beyond those of traditional desktop displays. Instead of rendering and warping…

Graphics · Computer Science 2018-06-15 Tobias Ritschel , Sebastian Friston , Anthony Steed

Latent representation learning has been an active field of study for decades in numerous applications. Inspired among others by the tokenization from Natural Language Processing and motivated by the research of a simple data representation,…

Signal Processing · Electrical Eng. & Systems 2024-09-26 Benoît Giniès , Xiaoyu Bie , Olivier Fercoq , Gaël Richard

We experimentally demonstrate the use of a novel liquid state, whispering-gallery-mode optical resonator as a highly sensitive humidity sensor. The optical resonator used consists of a droplet made of glycerol, a transparent liquid that…

Nearly all 3D displays need calibration for correct rendering. More often than not, the optical elements in a 3D display are misaligned from the designed parameter setting. As a result, 3D magic does not perform well as intended. The…

Computer Vision and Pattern Recognition · Computer Science 2017-04-26 Hyoseok Hwang , Hyun Sung Chang , Dongkyung Nam , In So Kweon

Current vision systems are trained on huge datasets, and these datasets come with costs: curation is expensive, they inherit human biases, and there are concerns over privacy and usage rights. To counter these costs, interest has surged in…

Computer Vision and Pattern Recognition · Computer Science 2022-05-02 Manel Baradad , Jonas Wulff , Tongzhou Wang , Phillip Isola , Antonio Torralba

Foley sound generation, the art of creating audio for multimedia, has recently seen notable advancements through text-conditioned latent diffusion models. These systems use multimodal text-audio representation models, such as Contrastive…

Audio and Speech Processing · Electrical Eng. & Systems 2024-10-15 Tornike Karchkhadze , Hassan Salami Kavaki , Mohammad Rasool Izadi , Bryce Irvin , Mikolaj Kegler , Ari Hertz , Shuo Zhang , Marko Stamenovic

Event cameras are a new type of sensors that are different from traditional cameras. Each pixel is triggered asynchronously by event. The trigger event is the change of the brightness irradiated on the pixel. If the increment or decrement…

Computer Vision and Pattern Recognition · Computer Science 2022-03-24 Kun Xiao , Guohui Wang , Yi Chen , Jinghong Nan , Yongfeng Xie

The safety of children in children home has become an increasing social concern, and the purpose of this experiment is to use machine learning applied to detect the scenarios of child abuse to increase the safety of children. This…

Audio and Speech Processing · Electrical Eng. & Systems 2023-07-31 Jiuqi Yan , Yingxian Chen , W. W. T. Fok

Speaker recognition is a biometric modality that uses underlying speech information to determine the identity of the speaker. Speaker Identification (SID) under noisy conditions is one of the challenging topics in the field of speech…

Sound · Computer Science 2019-08-02 Nursadul Mamun , Ria Ghosh , John H. L. Hansen

Realistic simulation is critical for applications ranging from robotics to animation. Traditional analytic simulators sometimes struggle to capture sufficiently realistic simulation which can lead to problems including the well known…

Acoustic levitation provides a unique method for manipulating small particles as it completely evades effects from gravity, container walls, or physical handling. These advantages make it a tantalizing platform for studying complex…

Soft Condensed Matter · Physics 2025-12-01 Sue Shi , Maximilian C. Hübl , Galien P. Grosjean , Carl P. Goodrich , Scott R. Waitukaitis
‹ Prev 1 8 9 10 Next ›