English
Related papers

Related papers: A Statistics-Driven Differentiable Approach for So…

200 papers

We present a system for automatic multi-axis perceptual quality prediction of generative audio, developed for Track 2 of the AudioMOS Challenge 2025. The task is to predict four Audio Aesthetic Scores--Production Quality, Production…

Audio and Speech Processing · Electrical Eng. & Systems 2025-09-04 Dyah A. M. G. Wisnu , Ryandhimas E. Zezario , Stefano Rini , Hsin-Min Wang , Yu Tsao

We present a novel texture synthesis framework, enabling the generation of infinite, high-quality 3D textures given a 2D exemplar image. Inspired by recent advances in natural texture synthesis, we train deep neural models to generate…

Computer Vision and Pattern Recognition · Computer Science 2020-07-01 Tiziano Portenier , Siavash Bigdeli , Orcun Goksel

We introduce TM-NET, a novel deep generative model for synthesizing textured meshes in a part-aware manner. Once trained, the network can generate novel textured meshes from scratch or predict textures for a given 3D mesh, without image…

Graphics · Computer Science 2021-06-10 Lin Gao , Tong Wu , Yu-Jie Yuan , Ming-Xian Lin , Yu-Kun Lai , Hao Zhang

Capturing the essence of a textile image in a robust way is important to retrieve it in a large repository, especially if it has been acquired in the wild (by taking a photo of the textile of interest). In this paper we show that a…

Computer Vision and Pattern Recognition · Computer Science 2019-10-07 Christian Joppi , Marco Godi , Andrea Giachetti , Fabio Pellacini , Marco Cristani

Despite recent progress in large-scale sound event detection (SED) systems capable of handling hundreds of sound classes, existing multi-class classification frameworks remain fundamentally limited. They cannot process free-text sound…

Audio and Speech Processing · Electrical Eng. & Systems 2025-09-24 Jiarui Hai , Helin Wang , Weizhe Guo , Mounya Elhilali

Generative models for speech synthesis face a fundamental trade-off: discrete tokens ensure stability but sacrifice expressivity, while continuous signals retain acoustic richness but suffer from error accumulation due to task entanglement.…

Sound content creation, essential for multimedia works such as video games and films, often involves extensive trial-and-error, enabling creators to semantically reflect their artistic ideas and inspirations, which evolve throughout the…

Current speech generation research can be categorized into two primary classes: non-autoregressive and autoregressive. The fundamental distinction between these approaches lies in the duration prediction strategy employed for…

Audio and Speech Processing · Electrical Eng. & Systems 2025-05-16 Linhan Ma , Dake Guo , He Wang , Jin Xu , Lei Xie

This paper presents a significant improvement for the synthesis of texture images using convolutional neural networks (CNNs), making use of constraints on the Fourier spectrum of the results. More precisely, the texture synthesis is…

Computer Vision and Pattern Recognition · Computer Science 2016-05-20 Gang Liu , Yann Gousseau , Gui-Song Xia

The influence of textures on machine learning models has been an ongoing investigation, specifically in texture bias/learning, interpretability, and robustness. However, due to the lack of large and diverse texture data available, the…

Computer Vision and Pattern Recognition · Computer Science 2025-05-13 Blaine Hoak , Patrick McDaniel

Detecting synthetic from real speech is increasingly crucial due to the risks of misinformation and identity impersonation. While various datasets for synthetic speech analysis have been developed, they often focus on specific areas,…

Sound · Computer Science 2025-07-18 Zhoulin Ji , Chenhao Lin , Hang Wang , Chao Shen

3D meshes are widely used in computer vision and graphics for their efficiency in animation and minimal memory use, playing a crucial role in movies, games, AR, and VR. However, creating temporally consistent and realistic textures for mesh…

Computer Vision and Pattern Recognition · Computer Science 2025-05-06 Jingzhi Bao , Xueting Li , Ming-Hsuan Yang

Voice Timbre Attribute Detection (vTAD) plays a pivotal role in fine-grained timbre modeling for speech generation tasks. However, it remains challenging due to the inherently subjective nature of timbre descriptors and the severe label…

Sound · Computer Science 2025-08-25 Zhiyu Wu , Jingyi Fang , Yufei Tang , Yuanzhong Zheng , Yaoxuan Wang , Haojun Fei

Existing knowledge distillation works for semantic segmentation mainly focus on transferring high-level contextual knowledge from teacher to student. However, low-level texture knowledge is also of vital importance for characterizing the…

Computer Vision and Pattern Recognition · Computer Science 2025-03-18 Deyi Ji , Haoran Wang , Mingyuan Tao , Jianqiang Huang , Xian-Sheng Hua , Hongtao Lu

Speech audio in the wild is often processed by post-production effects, but existing speech datasets rarely provide precise annotations of effects and parameters, limiting systematic study. We introduce VoxEffects, a speech audio effects…

Audio and Speech Processing · Electrical Eng. & Systems 2026-04-15 Zhe Zhang , Yigitcan Özer , Junichi Yamagishi

Procedural textures are normally generated from mathematical models with parameters carefully selected by experienced users. However, for naive users, the intuitive way to obtain a desired texture is to provide semantic descriptions such as…

Computer Vision and Pattern Recognition · Computer Science 2017-04-14 Junyu Dong , Lina Wang , Jun Liu , Xin Sun

Recent advancements in neural audio codecs have enabled the use of tokenized audio representations in various audio generation tasks, such as text-to-speech, text-to-audio, and text-to-music generation. Leveraging this approach, we propose…

Sound · Computer Science 2025-02-14 Kyungsu Kim , Junghyun Koo , Sungho Lee , Haesun Joung , Kyogu Lee

The crystallographic texture is a key organization feature of many technical and biological materials. In these materials, especially hierarchically structured ones, the preferential alignment of the nano constituents is heavily influencing…

Generating sound effects with controllable variations is a challenging task, traditionally addressed using sophisticated physical models that require in-depth knowledge of signal processing parameters and algorithms. In the era of…

Sound · Computer Science 2024-12-30 Yunyi Liu , Craig Jin

Recent advances in zero-shot text-to-speech (TTS), driven by language models, diffusion models and masked generation, have achieved impressive naturalness in speech synthesis. Nevertheless, stability and fidelity remain key challenges,…

Sound · Computer Science 2025-10-24 Hualei Wang , Na Li , Chuke Wang , Shu Wu , Zhifeng Li , Dong Yu
‹ Prev 1 4 5 6 7 8 10 Next ›