English
Related papers

Related papers: Evaluating Text-to-Image and Text-to-Video Synthes…

200 papers

Devising indicative evaluation metrics for the image generation task remains an open problem. The most widely used metric for measuring the similarity between real and generated images has been the Fr\'echet Inception Distance (FID) score.…

Computer Vision and Pattern Recognition · Computer Science 2020-06-30 Muhammad Ferjad Naeem , Seong Joon Oh , Youngjung Uh , Yunjey Choi , Jaejun Yoo

The growing popularity of generative music models underlines the need for perceptually relevant, objective music quality metrics. The Frechet Audio Distance (FAD) is commonly used for this purpose even though its correlation with perceptual…

Audio and Speech Processing · Electrical Eng. & Systems 2024-03-07 Azalea Gui , Hannes Gamper , Sebastian Braun , Dimitra Emmanouilidou

Fr\'echet Audio Distance (FAD) is the de facto standard for evaluating text-to-audio generation, yet its scores depend on the underlying encoder's embedding space. An encoder's training task dictates which acoustic features are preserved or…

Audio and Speech Processing · Electrical Eng. & Systems 2026-03-02 Wonwoo Jeong

Customization of text-to-image models enables users to insert new concepts or objects and generate them in unseen settings. Existing methods either rely on comparatively expensive test-time optimization or train encoders on single-image…

Computer Vision and Pattern Recognition · Computer Science 2025-10-14 Nupur Kumari , Xi Yin , Jun-Yan Zhu , Ishan Misra , Samaneh Azadi

Recently, open-domain text-to-video (T2V) generation models have made remarkable progress. However, the promising results are mainly shown by the qualitative cases of generated videos, while the quantitative evaluation of T2V models still…

Computer Vision and Pattern Recognition · Computer Science 2023-12-27 Yuanxin Liu , Lei Li , Shuhuai Ren , Rundong Gao , Shicheng Li , Sishuo Chen , Xu Sun , Lu Hou

Text-to-3D (T23D) generation has emerged as a crucial visual generation task, aiming at synthesizing 3D content from textual descriptions. Studies of this task are currently shifting from per-scene T23D, which requires optimization of the…

Computer Vision and Pattern Recognition · Computer Science 2025-12-04 Xiao Cai , Sitong Su , Jingkuan Song , Pengpeng Zeng , Ji Zhang , Qinhong Du , Mengqi Li , Heng Tao Shen , Lianli Gao

Language models are increasingly being incorporated as components in larger AI systems for various purposes, from prompt optimization to automatic evaluation. In this work, we analyze the construct validity of four recent, commonly used…

Computation and Language · Computer Science 2024-12-19 Candace Ross , Melissa Hall , Adriana Romero Soriano , Adina Williams

In this paper we introduce the Frechet Music Distance (FMD), a novel evaluation metric for generative symbolic music models, inspired by the Frechet Inception Distance (FID) in computer vision and Frechet Audio Distance (FAD) in generative…

Sound · Computer Science 2025-01-17 Jan Retkowski , Jakub Stępniak , Mateusz Modrzejewski

Image quality evaluation accurately is vital in developing image stitching algorithms as it directly reflects the algorithms progress. However, commonly used objective indicators always produce inconsistent and even conflicting results with…

Image and Video Processing · Electrical Eng. & Systems 2024-04-23 Xinrui Zhang , Shengwei Guo , Guobing Sun

Objective evaluation of synthetic speech quality remains a critical challenge. Human listening tests are the gold standard, but costly and impractical at scale. Fr\'echet Distance has emerged as a promising alternative, yet its reliability…

Sound · Computer Science 2026-01-30 June-Woo Kim , Dhruv Agarwal , Federica Cerina

Recent advancements in text-to-video models such as Sora, Gen-3, MovieGen, and CogVideoX are pushing the boundaries of synthetic video generation, with adoption seen in fields like robotics, autonomous driving, and entertainment. As these…

Computer Vision and Pattern Recognition · Computer Science 2025-04-28 S P Sharan , Minkyu Choi , Sahil Shah , Harsh Goel , Mohammad Omama , Sandeep Chinchali

The text-to-image synthesis by diffusion models has recently shown remarkable performance in generating high-quality images. Although performs well for simple texts, the models may get confused when faced with complex texts that contain…

Computer Vision and Pattern Recognition · Computer Science 2024-01-15 Chang Yu , Junran Peng , Xiangyu Zhu , Zhaoxiang Zhang , Qi Tian , Zhen Lei

Automatically describing an image with a sentence is a long-standing challenge in computer vision and natural language processing. Due to recent progress in object detection, attribute classification, action recognition, etc., there is…

Computer Vision and Pattern Recognition · Computer Science 2015-06-04 Ramakrishna Vedantam , C. Lawrence Zitnick , Devi Parikh

With advances in the quality of text-to-image (T2I) models has come interest in benchmarking their prompt faithfulness -- the semantic coherence of generated images to the prompts they were conditioned on. A variety of T2I faithfulness…

Computer Vision and Pattern Recognition · Computer Science 2024-11-01 Michael Saxon , Fatima Jahara , Mahsa Khoshnoodi , Yujie Lu , Aditya Sharma , William Yang Wang

Evaluating and comparing text-to-image models is a challenging problem. Significant advances in the field have recently been made, piquing interest of various industrial sectors. As a consequence, a gold standard in the field should cover a…

Computer Vision and Pattern Recognition · Computer Science 2022-12-16 Federico A. Galatolo , Mario G. C. A. Cimino , Edoardo Cogotti

Diffusion models (DMs) have recently gained attention with state-of-the-art performance in text-to-image synthesis. Abiding by the tradition in deep learning, DMs are trained and evaluated on the images with fixed sizes. However, users are…

Computer Vision and Pattern Recognition · Computer Science 2023-10-27 Zhiyu Jin , Xuli Shen , Bin Li , Xiangyang Xue

Evaluating the quality of videos generated from text-to-video (T2V) models is important if they are to produce plausible outputs that convince a viewer of their authenticity. We examine some of the metrics used in this area and highlight…

Computer Vision and Pattern Recognition · Computer Science 2023-09-18 Iya Chivileva , Philip Lynch , Tomas E. Ward , Alan F. Smeaton

Text-image generation has advanced rapidly, but assessing whether outputs truly capture the objects, attributes, and relations described in prompts remains a central challenge. Evaluation in this space relies heavily on automated metrics,…

Computer Vision and Pattern Recognition · Computer Science 2025-11-11 Seyed Amir Kasaei , Ali Aghayari , Arash Marioriyad , Niki Sepasian , MohammadAmin Fazli , Mahdieh Soleymani Baghshah , Mohammad Hossein Rohban

While text-conditional 3D object generation and manipulation have seen rapid progress, the evaluation of coherence between generated 3D shapes and input textual descriptions lacks a clear benchmark. The reason is twofold: a) the low quality…

Computer Vision and Pattern Recognition · Computer Science 2023-09-15 Andrea Amaduzzi , Giuseppe Lisanti , Samuele Salti , Luigi Di Stefano

Mitigating biases in machine learning models has become an increasing concern in Natural Language Processing (NLP), particularly in developing fair text embeddings, which are crucial yet challenging for real-world applications like search…

Computation and Language · Computer Science 2024-06-25 Wenlong Deng , Blair Chen , Beidi Zhao , Chiyu Zhang , Xiaoxiao Li , Christos Thrampoulidis