English
Related papers

Related papers: Evaluation of Automatic Video Captioning Using Dir…

200 papers

Recently, the state-of-the-art models for image captioning have overtaken human performance based on the most popular metrics, such as BLEU, METEOR, ROUGE, and CIDEr. Does this mean we have solved the task of image captioning? The above…

Computer Vision and Pattern Recognition · Computer Science 2019-05-16 Qingzhong Wang , Antoni B. Chan

Model-based evaluation metrics (e.g., CLIPScore and GPTScore) have demonstrated decent correlations with human judgments in various language generation tasks. However, their impact on fairness remains largely unexplored. It is widely…

Computation and Language · Computer Science 2023-11-06 Haoyi Qiu , Zi-Yi Dou , Tianlu Wang , Asli Celikyilmaz , Nanyun Peng

Automatic image aesthetics assessment is a computer vision problem dealing with categorizing images into different aesthetic levels. The categorization is usually done by analyzing an input image and computing some measure of the degree to…

Computer Vision and Pattern Recognition · Computer Science 2022-02-08 Abbas Anwar , Saira Kanwal , Muhammad Tahir , Muhammad Saqib , Muhammad Uzair , Mohammad Khalid Imam Rahmani , Habib Ullah

With the rapid advancement of text-conditioned Video Generation Models (VGMs), the quality of generated videos has significantly improved, bringing these models closer to functioning as ``*world simulators*'' and making real-world-level…

Artificial Intelligence · Computer Science 2025-04-22 Haotong Yang , Qingyuan Zheng , Yunjian Gao , Yongkun Yang , Yangbo He , Zhouchen Lin , Muhan Zhang

Video-language models (VLMs) learn to reason about the dynamic visual world through natural language. We introduce a suite of open datasets, benchmarks, and recipes for scalable oversight that enable precise video captioning. First, we…

Computer Vision and Pattern Recognition · Computer Science 2026-04-28 Zhiqiu Lin , Chancharik Mitra , Siyuan Cen , Isaac Li , Yuhan Huang , Yu Tong Tiffany Ling , Hewei Wang , Irene Pi , Shihang Zhu , Ryan Rao , George Liu , Jiaxi Li , Ruojin Li , Yili Han , Yilun Du , Deva Ramanan

Recent great advances in video generation models have demonstrated their potential to produce high-quality videos, bringing challenges to effective evaluation. Unlike human evaluation, existing automated evaluation metrics lack highlevel…

Computer Vision and Pattern Recognition · Computer Science 2025-12-23 Zhun Mou , Bin Xia , Zhengchao Huang , Wenming Yang , Jiaya Jia

We propose VC-Inspector, a lightweight, open-source large multimodal model (LMM) for reference-free evaluation of video captions, with a focus on factual accuracy. Unlike existing metrics that suffer from limited context handling, weak…

Computer Vision and Pattern Recognition · Computer Science 2026-04-21 Shubhashis Roy Dipta , Tz-Ying Wu , Subarna Tripathi

We present a method that automatically evaluates emotional response from spontaneous facial activity recorded by a depth camera. The automatic evaluation of emotional response, or affect, is a fascinating challenge with many applications,…

Human-Computer Interaction · Computer Science 2017-01-20 Daniel Hadar

Evaluation metrics for image captioning face two challenges. Firstly, commonly used metrics such as CIDEr, METEOR, ROUGE and BLEU often do not correlate well with human judgments. Secondly, each metric has well known blind spots to…

Computer Vision and Pattern Recognition · Computer Science 2018-06-19 Yin Cui , Guandao Yang , Andreas Veit , Xun Huang , Serge Belongie

Accelerated by the tremendous increase in Internet bandwidth and storage space, video data has been generated, published and spread explosively, becoming an indispensable part of today's big data. In this paper, we focus on reviewing two…

Computer Vision and Pattern Recognition · Computer Science 2018-02-23 Zuxuan Wu , Ting Yao , Yanwei Fu , Yu-Gang Jiang

The emergence of contemporary deepfakes has attracted significant attention in machine learning research, as artificial intelligence (AI) generated synthetic media increases the incidence of misinterpretation and is difficult to distinguish…

Computer Vision and Pattern Recognition · Computer Science 2024-11-12 Ammarah Hashmi , Sahibzada Adil Shahzad , Chia-Wen Lin , Yu Tsao , Hsin-Min Wang

The Automated Audio Captioning (AAC) task asks models to generate natural language descriptions of an audio input. Evaluating these machine-generated audio captions is a complex task that requires considering diverse factors, among them,…

Computation and Language · Computer Science 2025-08-12 Tsung-Han Wu , Joseph E. Gonzalez , Trevor Darrell , David M. Chan

Automatic description generation from natural images is a challenging problem that has recently received a large amount of interest from the computer vision and natural language processing communities. In this survey, we classify the…

Generative models have made immense progress in recent years, particularly in their ability to generate high quality images. However, that quality has been difficult to evaluate rigorously, with evaluation dominated by heuristic approaches…

Computer Vision and Pattern Recognition · Computer Science 2019-12-30 Y. Alex Kolchinski , Sharon Zhou , Shengjia Zhao , Mitchell Gordon , Stefano Ermon

This paper focuses on enhancing the captions generated by image-caption generation systems. We propose an approach for improving caption generation systems by choosing the most closely related output to the image rather than the most likely…

Computation and Language · Computer Science 2023-07-10 Ahmed Sabir

Image captioning as a multimodal task has drawn much interest in recent years. However, evaluation for this task remains a challenging problem. Existing evaluation metrics focus on surface similarity between a candidate caption and a set of…

Computation and Language · Computer Science 2019-12-20 Huiyuan Xie , Tom Sherborne , Alexander Kuhnle , Ann Copestake

This paper presents a novel approach to perform sentiment analysis of news videos, based on the fusion of audio, textual and visual clues extracted from their contents. The proposed approach aims at contributing to the semiodiscoursive…

Computation and Language · Computer Science 2016-04-12 Moisés H. R. Pereira , Flávio L. C. Pádua , Adriano C. M. Pereira , Fabrício Benevenuto , Daniel H. Dalip

Beyond conventional paradigms of translating speech and text, recently, there has been interest in automated transcreation of images to facilitate localization of visual content across different cultures. Attempts to define this as a formal…

Computation and Language · Computer Science 2025-03-24 Simran Khanuja , Vivek Iyer , Claire He , Graham Neubig

Image captioning studies heavily rely on automatic evaluation metrics such as BLEU and METEOR. However, such n-gram-based metrics have been shown to correlate poorly with human evaluation, leading to the proposal of alternative metrics such…

Computer Vision and Pattern Recognition · Computer Science 2023-11-08 Yuiga Wada , Kanta Kaneda , Komei Sugiura

Evaluating image captions typically relies on reference captions, which are costly to obtain and exhibit significant diversity and subjectivity. While reference-free evaluation metrics have been proposed, most focus on cross-modal…

Computer Vision and Pattern Recognition · Computer Science 2025-01-09 Tianyu Cui , Jinbin Bai , Guo-Hua Wang , Qing-Guo Chen , Zhao Xu , Weihua Luo , Kaifu Zhang , Ye Shi
‹ Prev 1 4 5 6 7 8 10 Next ›