English
Related papers

Related papers: Challenge on Sound Scene Synthesis: Evaluating Tex…

200 papers

Can machines recording an audio-visual scene produce realistic, matching audio-visual experiences at novel positions and novel view directions? We answer it by studying a new task -- real-world audio-visual scene synthesis -- and a…

Computer Vision and Pattern Recognition · Computer Science 2023-10-17 Susan Liang , Chao Huang , Yapeng Tian , Anurag Kumar , Chenliang Xu

Humans can imagine a scene from a sound. We want machines to do so by using conditional generative adversarial networks (GANs). By applying the techniques including spectral norm, projection discriminator and auxiliary classifier, compared…

Computation and Language · Computer Science 2018-08-14 Chia-Hung Wan , Shun-Po Chuang , Hung-Yi Lee

Advances in AI-generated content have led to wide adoption of large language models, diffusion-based visual generators, and synthetic audio tools. However, these developments raise critical concerns about misinformation, copyright…

Computation and Language · Computer Science 2025-09-30 Lele Cao

Text-to-audio (TTA) generation is advancing rapidly, but evaluation remains challenging because human listening studies are expensive and existing automatic metrics capture only limited aspects of perceptual quality. We introduce AudioEval,…

Sound · Computer Science 2026-01-30 Hui Wang , Jinghua Zhao , Junyang Cheng , Cheng Liu , Yuhang Jia , Haoqin Sun , Jiaming Zhou , Yong Qin

In this work, we investigate the personalization of text-to-music diffusion models in a few-shot setting. Motivated by recent advances in the computer vision domain, we are the first to explore the combination of pre-trained text-to-audio…

In recent years, image generation has shown a great leap in performance, where diffusion models play a central role. Although generating high-quality images, such models are mainly conditioned on textual descriptions. This begs the…

Sound · Computer Science 2023-05-23 Guy Yariv , Itai Gat , Lior Wolf , Yossi Adi , Idan Schwartz

Recent advancements in 4D generation have demonstrated its remarkable capability in synthesizing photorealistic renderings of dynamic 3D scenes. However, despite achieving impressive visual performance, almost all existing methods overlook…

Sound · Computer Science 2026-03-02 Siyi Xie , Hanxin Zhu , Xinyi Chen , Tianyu He , Xin Li , Zhibo Chen

Audio textures are a subset of environmental sounds, often defined as having stable statistical characteristics within an adequately large window of time but may be unstructured locally. They include common everyday sounds such as from…

Sound · Computer Science 2020-11-26 M. Huzaifah , L. Wyse

Audio editing aims to manipulate audio content based on textual descriptions, supporting tasks such as adding, removing, or replacing audio events. Despite recent progress, the lack of high-quality benchmark datasets and comprehensive…

Sound · Computer Science 2026-02-03 Yuhang Jia , Hui Wang , Xin Nie , Yujie Guo , Lianru Gao , Yong Qin

The Inspirational and Convincing Audio Generation Challenge 2024 (ICAGC 2024) is part of the ISCSLP 2024 Competitions and Challenges track. While current text-to-speech (TTS) technology can generate high-quality audio, its ability to convey…

Audio and Speech Processing · Electrical Eng. & Systems 2024-08-01 Ruibo Fu , Rui Liu , Chunyu Qiang , Yingming Gao , Yi Lu , Shuchen Shi , Tao Wang , Ya Li , Zhengqi Wen , Chen Zhang , Hui Bu , Yukun Liu , Xin Qi , Guanjun Li

We introduce a novel pipeline for joint audio-visual editing that enhances the coherence between edited video and its accompanying audio. Our approach first applies state-of-the-art video editing techniques to produce the target video, then…

Multimedia · Computer Science 2026-03-18 Masato Ishii , Akio Hayakawa , Takashi Shibuya , Yuki Mitsufuji

The requirement of large amounts of annotated images has become one grand challenge while training deep neural network models for various visual detection and recognition tasks. This paper presents a novel image synthesis technique that…

Computer Vision and Pattern Recognition · Computer Science 2018-09-27 Fangneng Zhan , Shijian Lu , Chuhui Xue

Recent progress in diffusion-based audio generation and restoration has substantially improved performance across heterogeneous conditioning regimes, including text-conditioned audio generation and audio-conditioned super-resolution.…

Sound · Computer Science 2026-05-07 Xuanhao Zhang , Chang Li

Recent advances in video generation have been remarkable, enabling models to produce visually compelling videos with synchronized audio. While existing video generation benchmarks provide comprehensive metrics for visual quality, they lack…

Computer Vision and Pattern Recognition · Computer Science 2026-04-07 Daili Hua , Xizhi Wang , Bohan Zeng , Xinyi Huang , Hao Liang , Junbo Niu , Xinlong Chen , Quanqing Xu , Wentao Zhang

Text-to-video (T2V) generative models have advanced significantly, yet their ability to compose different objects, attributes, actions, and motions into a video remains unexplored. Previous text-to-video benchmarks also neglect this…

Computer Vision and Pattern Recognition · Computer Science 2025-01-16 Kaiyue Sun , Kaiyi Huang , Xian Liu , Yue Wu , Zihan Xu , Zhenguo Li , Xihui Liu

AI-based text-to-image models do not only excel at generating realistic images, they also give designers more and more fine-grained control over the image content. Consequently, these approaches have gathered increased attention within the…

Computer Vision and Pattern Recognition · Computer Science 2025-01-30 Sebastian Hartwig , Dominik Engel , Leon Sick , Hannah Kniesel , Tristan Payer , Poonam Poonam , Michael Glöckler , Alex Bäuerle , Timo Ropinski

Many recently published Text-to-Speech (TTS) systems produce audio close to real speech. However, TTS evaluation needs to be revisited to make sense of the results obtained with the new architectures, approaches and datasets. We propose…

Audio and Speech Processing · Electrical Eng. & Systems 2024-12-03 Christoph Minixhofer , Ondřej Klejch , Peter Bell

Text to speech (TTS), or speech synthesis, which aims to synthesize intelligible and natural speech given text, is a hot research topic in speech, language, and machine learning communities and has broad applications in the industry. As the…

Audio and Speech Processing · Electrical Eng. & Systems 2021-07-26 Xu Tan , Tao Qin , Frank Soong , Tie-Yan Liu

Evaluating text-to-image and text-to-video models is challenging due to a fundamental disconnect: established metrics fail to jointly measure visual quality and semantic alignment with text, leading to a poor correlation with human…

Computer Vision and Pattern Recognition · Computer Science 2025-08-28 Jaywon Koo , Jefferson Hernandez , Moayed Haji-Ali , Ziyan Yang , Vicente Ordonez

We focus on the foundational task of Scene Staging: given a reference scene image and a text condition specifying an actor category to be generated in the scene and its spatial relation to the scene, the goal is to synthesize an output…

Computer Vision and Pattern Recognition · Computer Science 2026-05-19 Cong Xie , Che Wang , Yan Zhang , Ruiqi Yu , Han Zou , Zheng Pan , Zhenpeng Zhan
‹ Prev 1 3 4 5 6 7 10 Next ›