English
Related papers

Related papers: T2A-Feedback: Improving Basic Capabilities of Text…

200 papers

Rapid advances in audio-video (AV) generation have enabled high-fidelity synthesis with synchronized sound, particularly for human-related scenarios involving speech and interactions. Yet evaluation for AV generation remains at an early…

Artificial Intelligence · Computer Science 2026-05-26 Jialiang Yang , Bin Xia , Ruihang Chu , Dingdong Wang , Wanke Xia , Zhun Mou , Tianyang Zhong , Yiting Zhao , Wenming Yang

Automated audio captioning (AAC) aims to generate informative descriptions for various sounds from nature and/or human activities. In recent years, AAC has quickly attracted research interest, with state-of-the-art systems now relying on a…

Generative Artificial Intelligence (AI) has enabled the development of sophisticated models that are capable of producing high-caliber text, images, and other outputs through the utilization of large pre-trained models. Nevertheless,…

Computation and Language · Computer Science 2023-02-14 Jinlan Fu , See-Kiong Ng , Zhengbao Jiang , Pengfei Liu

The rapid evolution of end-to-end spoken dialogue systems demands transcending mere textual semantics to incorporate paralinguistic nuances and the spontaneous nature of human conversation. However, current methods struggle with two…

Audio and Speech Processing · Electrical Eng. & Systems 2026-05-12 Jingyu Lu , Yuhan Wang , Fan Zhuo , Xize Cheng , Changhao Pan , Xueyi Pu , Yifu Chen , Chenyuhao Wen , Tianle Liang , Zhou Zhao

Significant advancements in video diffusion models have brought substantial progress to the field of text-to-video (T2V) synthesis. However, existing T2V synthesis model struggle to accurately generate complex motion dynamics, leading to a…

Computer Vision and Pattern Recognition · Computer Science 2024-06-06 Haoran Cheng , Liang Peng , Linxuan Xia , Yuepeng Hu , Hengjia Li , Qinglin Lu , Xiaofei He , Boxi Wu

Controlled text generation is a very important task in the arena of natural language processing due to its promising applications. In order to achieve this task we mainly introduce the novel soft prompt tuning method of using soft prompts…

Computation and Language · Computer Science 2022-12-07 Damith Chamalke Senadeera , Julia Ive

The abilities of Generative-Artificial Intelligence (AI) to produce real-time, sophisticated responses across diverse contexts has promised a huge potential in physics education, particularly in providing customized feedback. In this study,…

Physics Education · Physics 2025-08-14 Amogh Sirnoorkar , N. Sanjay Rebello

Text-to-image (T2I) models achieve high-fidelity generation through extensive training on large datasets. However, these models may unintentionally pick up undesirable biases of their training data, such as over-representation of particular…

Computer Vision and Pattern Recognition · Computer Science 2024-07-01 Shufan Li , Harkanwar Singh , Aditya Grover

Recent progress in audio-language modeling, such as automated audio captioning, has benefited from training on synthetic data generated with the aid of large-language models. However, such approaches for environmental sound captioning have…

Sound · Computer Science 2024-10-17 Mithun Manivannan , Vignesh Nethrapalli , Mark Cartwright

The field of text-to-image (T2I) generation has garnered significant attention both within the research community and among everyday users. Despite the advancements of T2I models, a common issue encountered by users is the need for…

Computation and Language · Computer Science 2023-10-31 Wanrong Zhu , Xinyi Wang , Yujie Lu , Tsu-Jui Fu , Xin Eric Wang , Miguel Eckstein , William Yang Wang

Diffusion models have emerged as powerful deep generative techniques, producing high-quality and diverse samples in applications in various domains including audio. While existing reviews provide overviews, there remains limited in-depth…

Sound · Computer Science 2026-01-16 Ge Zhu , Yutong Wen , Zhiyao Duan

In recent years, fueled by the rapid advancement of diffusion models, text-to-video (T2V) generation models have achieved remarkable progress, with notable examples including Pika, Luma, Kling, and Open-Sora. Although these models exhibit…

Computer Vision and Pattern Recognition · Computer Science 2025-06-18 Jiayang Liu , Siyuan Liang , Shiqian Zhao , Rongcheng Tu , Wenbo Zhou , Aishan Liu , Dacheng Tao , Siew Kei Lam

Existing dominant methods for audio generation include Generative Adversarial Networks (GANs) and diffusion-based methods like Flow Matching. GANs suffer from slow convergence during training, while diffusion methods require multi-step…

Audio and Speech Processing · Electrical Eng. & Systems 2026-03-10 Zengwei Yao , Wei Kang , Han Zhu , Liyong Guo , Lingxuan Ye , Fangjun Kuang , Weiji Zhuang , Zhaoqing Li , Zhifeng Han , Long Lin , Daniel Povey

We present Synthio, a novel approach for augmenting small-scale audio classification datasets with synthetic data. Our goal is to improve audio classification accuracy with limited labeled data. Traditional data augmentation techniques,…

Audio and Speech Processing · Electrical Eng. & Systems 2025-03-13 Sreyan Ghosh , Sonal Kumar , Zhifeng Kong , Rafael Valle , Bryan Catanzaro , Dinesh Manocha

Text-to-image (T2I) models have made substantial progress in generating images from textual prompts. However, they frequently fail to produce images consistent with physical commonsense, a vital capability for applications in world…

Computer Vision and Pattern Recognition · Computer Science 2024-09-24 Fanqing Meng , Wenqi Shao , Lixin Luo , Yahong Wang , Yiran Chen , Quanfeng Lu , Yue Yang , Tianshuo Yang , Kaipeng Zhang , Yu Qiao , Ping Luo

We present FastPitch, a fully-parallel text-to-speech model based on FastSpeech, conditioned on fundamental frequency contours. The model predicts pitch contours during inference. By altering these predictions, the generated speech can be…

Audio and Speech Processing · Electrical Eng. & Systems 2021-02-17 Adrian Łańcucki

The rapid development of diffusion models has significantly advanced AI-generated content (AIGC), particularly in Text-to-Image (T2I) and Text-to-Video (T2V) generation. Text-based video editing, leveraging these generative capabilities,…

Computer Vision and Pattern Recognition · Computer Science 2025-01-22 Yupeng Chen , Penglin Chen , Xiaoyu Zhang , Yixian Huang , Qian Xie

Text-to-image (T2I) models today are capable of producing photorealistic, instruction-following images, yet they still frequently fail on prompts that require implicit world knowledge. Existing evaluation protocols either emphasize…

Computer Vision and Pattern Recognition · Computer Science 2025-12-12 Tianyang Han , Junhao Su , Junjie Hu , Peizhen Yang , Hengyu Shi , Junfeng Luo , Jialin Gao

Text-to-audio models are a type of generative model that produces audio output in response to a given textual prompt. Although level generators and the properties of the functional content that they create (e.g., playability) dominate most…

Sound · Computer Science 2025-11-17 Jonathan Morse , Azadeh Naderi , Swen Gaudl , Mark Cartwright , Amy K. Hoover , Mark J. Nelson

How does textual representation of audio relate to the Large Language Model's (LLMs) learning about the audio world? This research investigates the extent to which LLMs can be prompted to generate audio, despite their primary training in…

‹ Prev 1 8 9 10 Next ›