English
Related papers

Related papers: A Statistics-Driven Differentiable Approach for So…

200 papers

Visual AutoRegressive (VAR) models based on next-scale prediction enable efficient hierarchical generation, yet the inference cost grows quadratically at high resolutions. We observe that the computationally intensive later scales…

Computer Vision and Pattern Recognition · Computer Science 2026-03-03 Keli Liu , Zhendong Wang , Wengang Zhou , Houqiang Li

We present a deep neural network-based methodology for synthesising percussive sounds with control over high-level timbral characteristics of the sounds. This approach allows for intuitive control of a synthesizer, enabling the user to…

Audio and Speech Processing · Electrical Eng. & Systems 2020-04-06 António Ramires , Pritish Chandna , Xavier Favory , Emilia Gómez , Xavier Serra

This study introduces a novel and interpretable model, DiffVox, for matching vocal effects in music production. DiffVox, short for ``Differentiable Vocal Fx", integrates parametric equalisation, dynamic range control, delay, and reverb with…

Token-based language modeling is a prominent approach for speech generation, where tokens are obtained by quantizing features from self-supervised learning (SSL) models and extracting codes from neural speech codecs, generally referred to…

Audio and Speech Processing · Electrical Eng. & Systems 2025-06-30 Yang Yang , Yunpeng Li , George Sung , Shao-Fu Shih , Craig Dooley , Alessio Centazzo , Ramanan Rajeswaran

Recent neural network strategies for source separation attempt to model audio signals by processing their waveforms directly. Mean squared error (MSE) that measures the Euclidean distance between waveforms of denoised speech and the…

Audio and Speech Processing · Electrical Eng. & Systems 2018-06-05 Shrikant Venkataramani , Ryley Higa , Paris Smaragdis

Text-to-texture generation has recently attracted increasing attention, but existing methods often suffer from the problems of view inconsistencies, apparent seams, and misalignment between textures and the underlying mesh. In this paper,…

Computer Vision and Pattern Recognition · Computer Science 2024-10-01 Jangyeong Kim , Donggoo Kang , Junyoung Choi , Jeonga Wi , Junho Gwon , Jiun Bae , Dumim Yoon , Junghyun Han

Accurate prediction of perceptual attributes of haptic textures is essential for advancing VR and AR applications and enhancing robotic interaction with physical surfaces. This paper presents a deep learning-based multi-modal framework,…

Human-Computer Interaction · Computer Science 2025-06-24 Mudassir Ibrahim Awan , Seokhee Jeon

Diffusion models have recently advanced photorealistic human synthesis, although practical talking-head generation (THG) remains constrained by high inference latency, temporal instability such as flicker and identity drift, and imperfect…

Computer Vision and Pattern Recognition · Computer Science 2026-03-09 Soumya Mazumdar , Vineet Kumar Rakesh

Wafer defect segmentation is pivotal for semiconductor yield optimization yet remains challenged by the intrinsic conflict between microscale anomalies and highly periodic, overwhelming background textures. Existing deep learning paradigms…

Computer Vision and Pattern Recognition · Computer Science 2026-03-17 Zihan Zhang

In this paper, we present TexPro, a novel method for high-fidelity material generation for input 3D meshes given text prompts. Unlike existing text-conditioned texture generation methods that typically generate RGB textures with baked…

Graphics · Computer Science 2025-05-20 Ziqiang Dang , Wenqi Dong , Zesong Yang , Bangbang Yang , Liang Li , Yuewen Ma , Zhaopeng Cui

Global Style Tokens (GSTs) are a recently-proposed method to learn latent disentangled representations of high-dimensional data. GSTs can be used within Tacotron, a state-of-the-art end-to-end text-to-speech synthesis system, to uncover…

Computation and Language · Computer Science 2018-08-07 Daisy Stanton , Yuxuan Wang , RJ Skerry-Ryan

PESQ and POLQA , are standards are standards for automated assessment of voice quality of speech as experienced by human beings. The predictions of those objective measures should come as close as possible to subjective quality scores as…

Sound · Computer Science 2017-08-22 Dan Elbaz , Michael Zibulevsky

Recent attempts to interleave autoregressive (AR) sketchers with diffusion-based refiners over continuous speech representations have shown promise, but they remain brittle under distribution shift and offer limited levers for…

Audio and Speech Processing · Electrical Eng. & Systems 2025-10-17 Yakun Song , Xiaobin Zhuang , Jiawei Chen , Zhikang Niu , Guanrou Yang , Chenpeng Du , Dongya Jia , Zhuo Chen , Yuping Wang , Yuxuan Wang , Xie Chen

Conventional CNNs for texture synthesis consist of a sequence of (de)-convolution and up/down-sampling layers, where each layer operates locally and lacks the ability to capture the long-term structural dependency required by texture…

Computer Vision and Pattern Recognition · Computer Science 2020-07-15 Guilin Liu , Rohan Taori , Ting-Chun Wang , Zhiding Yu , Shiqiu Liu , Fitsum A. Reda , Karan Sapra , Andrew Tao , Bryan Catanzaro

Prevailing 3D texture generation methods, which often rely on multi-view fusion, are frequently hindered by inter-view inconsistencies and incomplete coverage of complex surfaces, limiting the fidelity and completeness of the generated…

Computer Vision and Pattern Recognition · Computer Science 2025-12-03 Yifei Zeng , Yajie Bao , Jiachen Qian , Shuang Wu , Youtian Lin , Hao Zhu , Buyu Li , Feihu Zhang , Xun Cao , Yao Yao

Accurately estimating and simulating the physical properties of objects from real-world sound recordings is of great practical importance in the fields of vision, graphics, and robotics. However, the progress in these directions has been…

Sound · Computer Science 2024-09-23 Xutong Jin , Chenxi Xu , Ruohan Gao , Jiajun Wu , Guoping Wang , Sheng Li

Despite the initial belief that Convolutional Neural Networks (CNNs) are driven by shapes to perform visual recognition tasks, recent evidence suggests that texture bias in CNNs provides higher performing models when learning on large…

Computer Vision and Pattern Recognition · Computer Science 2020-12-25 Reza Azad , Abdur R Fayjie , Claude Kauffman , Ismail Ben Ayed , Marco Pedersoli , Jose Dolz

This paper introduces StyleSpeech, a novel Text-to-Speech~(TTS) system that enhances the naturalness and accuracy of synthesized speech. Building upon existing TTS technologies, StyleSpeech incorporates a unique Style Decorator structure…

Sound · Computer Science 2024-12-31 Haowei Lou , Helen Paik , Wen Hu , Lina Yao

Texture is an important spatial feature which plays a vital role in content based image retrieval. The enormous growth of the internet and the wide use of digital data have increased the need for both efficient image database creation and…

Computer Vision and Pattern Recognition · Computer Science 2011-11-11 B. Vijayalakshmi , V. Subbiah Bharathi

Many audio processing tasks require perceptual assessment. The ``gold standard`` of obtaining human judgments is time-consuming, expensive, and cannot be used as an optimization criterion. On the other hand, automated metrics are efficient…

Audio and Speech Processing · Electrical Eng. & Systems 2020-05-19 Pranay Manocha , Adam Finkelstein , Richard Zhang , Nicholas J. Bryan , Gautham J. Mysore , Zeyu Jin