English
Related papers

Related papers: The Solution for the CVPR2024 NICE Image Captionin…

200 papers

In this paper, we introduce a model designed to improve the prediction of image-text alignment, targeting the challenge of compositional understanding in current visual-language models. Our approach focuses on generating high-quality…

Computer Vision and Pattern Recognition · Computer Science 2024-10-02 Yuheng Li , Haotian Liu , Mu Cai , Yijun Li , Eli Shechtman , Zhe Lin , Yong Jae Lee , Krishna Kumar Singh

Image captioning, an important vision-language task, often requires a tremendous number of finely labeled image-caption pairs for learning the underlying alignment between images and texts. In this paper, we proposed a multimodal data…

Computer Vision and Pattern Recognition · Computer Science 2023-11-14 Changrong Xiao , Sean Xin Xu , Kunpeng Zhang

Grounding-based vision and language models have been successfully applied to low-level vision tasks, aiming to precisely locate objects referred in captions. The effectiveness of grounding representation learning heavily relies on the scale…

Computer Vision and Pattern Recognition · Computer Science 2023-11-07 Jingru Yi , Burak Uzkent , Oana Ignat , Zili Li , Amanmeet Garg , Xiang Yu , Linda Liu

There are a thousand ways to caption an image. Contrastive Language Pretraining (CLIP) on the other hand, works by mapping an image and its caption to a single vector -- limiting how well CLIP-like models can represent the diverse ways to…

Computer Vision and Pattern Recognition · Computer Science 2025-04-01 Samuel Lavoie , Polina Kirichenko , Mark Ibrahim , Mahmoud Assran , Andrew Gordon Wilson , Aaron Courville , Nicolas Ballas

In scholarly documents, figures provide a straightforward way of communicating scientific findings to readers. Automating figure caption generation helps move model understandings of scientific documents beyond text and will help authors…

Computer Vision and Pattern Recognition · Computer Science 2023-06-07 Zhishen Yang , Raj Dabre , Hideki Tanaka , Naoaki Okazaki

Zero-shot Image Captioning (ZIC) increasingly utilizes synthetic datasets generated by text-to-image (T2I) models to mitigate the need for costly manual annotation. However, these T2I models often produce images that exhibit semantic…

Computer Vision and Pattern Recognition · Computer Science 2025-07-25 Si-Woo Kim , MinJu Jeon , Ye-Chan Kim , Soeun Lee , Taewhan Kim , Dong-Jin Kim

Supervised image captioning approaches have made great progress, but it is challenging to collect high-quality human-annotated image-text data. Recently, large-scale vision and language models (e.g., CLIP) and large-scale generative…

Computer Vision and Pattern Recognition · Computer Science 2024-03-29 Yiyu Wang , Hao Luo , Jungang Xu , Yingfei Sun , Fan Wang

Image captioning evaluation remains a significant challenge, as vision-language models evolve toward more challenging capabilities such as generating long-form and context-rich descriptions. State-of-the-art evaluation metrics involve…

Computer Vision and Pattern Recognition · Computer Science 2026-05-22 Gonçalo Gomes , Bruno Martins , Chrysoula Zerva

It is encouraged to see that progress has been made to bridge videos and natural language. However, mainstream video captioning methods suffer from slow inference speed due to the sequential manner of autoregressive decoding, and prefer…

Computer Vision and Pattern Recognition · Computer Science 2021-03-25 Bang Yang , Yuexian Zou , Fenglin Liu , Can Zhang

Image captioning, which generates natural language descriptions of the visual information in an image, is a crucial task in vision-language research. Previous models have typically addressed this task by aligning the generative capabilities…

Computer Vision and Pattern Recognition · Computer Science 2024-09-02 Qian Cao , Xu Chen , Ruihua Song , Xiting Wang , Xinting Huang , Yuchen Ren

Automatically captioning images with natural language sentences is an important research topic. State of the art models are able to produce human-like sentences. These models typically describe the depicted scene as a whole and do not…

Computer Vision and Pattern Recognition · Computer Science 2019-08-07 Philipp Harzig , Stephan Brehm , Rainer Lienhart , Carolin Kaiser , René Schallner

Data visualization captions help readers understand the purpose of a visualization and are crucial for individuals with visual impairments. The prevalence of poor figure captions and the successful application of deep learning approaches to…

Computer Vision and Pattern Recognition · Computer Science 2022-07-18 Anita Mahinpei , Zona Kostic , Chris Tanner

Automatically creating the description of an image using any natural languages sentence like English is a very challenging task. It requires expertise of both image processing as well as natural language processing. This paper discuss about…

Computer Vision and Pattern Recognition · Computer Science 2018-10-03 Parth Shah , Vishvajit Bakarola , Supriya Pati

Sequence-level learning objective has been widely used in captioning tasks to achieve the state-of-the-art performance for many models. In this objective, the model is trained by the reward on the quality of its generated captions…

Computer Vision and Pattern Recognition · Computer Science 2020-03-10 Jia Chen , Qin Jin

Significant performance gains in deep learning coupled with the exponential growth of image and video data on the Internet have resulted in the recent emergence of automated image captioning systems. Ensuring scalability of automated image…

Computer Vision and Pattern Recognition · Computer Science 2016-06-07 Karan Sharma , Arun CS Kumar , Suchendra Bhandarkar

This paper presents the NTIRE 2026 image super-resolution ($\times$4) challenge, one of the associated competitions of the NTIRE 2026 Workshop at CVPR 2026. The challenge aims to reconstruct high-resolution (HR) images from low-resolution…

Computer Vision and Pattern Recognition · Computer Science 2026-04-17 Zheng Chen , Kai Liu , Jingkai Wang , Xianglong Yan , Jianze Li , Ziqing Zhang , Jue Gong , Jiatong Li , Lei Sun , Xiaoyang Liu , Radu Timofte , Yulun Zhang , Jihye Park , Yoonjin Im , Hyungju Chun , Hyunhee Park , MinKyu Park , Zheng Xie , Xiangyu Kong , Weijun Yuan , Zhan Li , Qiurong Song , Luen Zhu , Fengkai Zhang , Xinzhe Zhu , Junyang Chen , Congyu Wang , Yixin Yang , Zhaorun Zhou , Jiangxin Dong , Jinshan Pan , Shengwei Wang , Jiajie Ou , Baiang Li , Sizhuo Ma , Qiang Gao , Jusheng Zhang , Jian Wang , Keze Wang , Yijiao Liu , Yingsi Chen , Hui Li , Yu Wang , Congchao Zhu , Saeed Ahmad , Ik Hyun Lee , Jun Young Park , Ji Hwan Yoon , Kainan Yan , Zian Wang , Weibo Wang , Shihao Zou , Chao Dong , Wei Zhou , Linfeng Li , Jaeseong Lee , Jaeho Chae , Jinwoo Kim , Seonjoo Kim , Yucong Hong , Zhenming Yan , Junye Chen , Ruize Han , Song Wang , Yuxuan Jiang , Chengxi Zeng , Tianhao Peng , Fan Zhang , David Bull , Tongyao Mu , Qiong Cao , Yifan Wang , Youwei Pan , Leilei Cao , Xiaoping Peng , Wei Deng , Yifei Chen , Wenbo Xiong , Xian Hu , Yuxin Zhang , Xiaoyun Cheng , Yang Ji , Zonghao Chen , Zhihao Xue , Junqin Hu , Nihal Kumar , Snehal Singh Tomar , Klaus Mueller , Surya Vashisth , Prateek Shaily , Jayant Kumar , Hardik Sharma , Ashish Negi , Sachin Chaudhary , Akshay Dudhane , Praful Hambarde , Amit Shukla , Shijun Shi , Jiangning Zhang , Yong Liu , Kai Hu , Jing Xu , Xianfang Zeng , Amitesh M , Hariharan S , Chia-Ming Lee , Yu-Fan Lin , Chih-Chung Hsu , Nishalini K , Sreenath K A , Bilel Benjdira , Anas M. Ali , Wadii Boulila , Shuling Zheng , Zhiheng Fu , Feng Zhang , Zhanglu Chen , Boyang Yao , Nikhil Pathak , Aagam Jain , Milan Kumar , Kishor Upla , Vivek Chavda , Sarang N S , Raghavendra Ramachandra , Zhipeng Zhang , Qi Wang , Shiyu Wang , Jiachen Tu , Guoyi Xu , Yaoxin Jiang , Jiajia Liu , Yaokun Shi , Yuqi Li , Chuanguang Yang , Weilun Feng , Zhuzhi Hong , Hao Wu , Junming Liu , Yingli Tian , Amish Bhushan Kulkarni , Tejas R R Shet , Saakshi M Vernekar , Nikhil Akalwadi , Kaushik Mallibhat , Ramesh Ashok Tabib , Uma Mudenagudi , Yuwen Pan , Tianrun Chen , Deyi Ji , Qi Zhu , Lanyun Zhu , Heyan Zhangyi

We explore an efficient approach to establish a foundational video-text model. We present VideoCoCa that maximally reuses a pretrained image-text contrastive captioner (CoCa) model and adapt it to video-text tasks with minimal extra…

Computer Vision and Pattern Recognition · Computer Science 2023-03-17 Shen Yan , Tao Zhu , Zirui Wang , Yuan Cao , Mi Zhang , Soham Ghosh , Yonghui Wu , Jiahui Yu

Automatically generating natural language descriptions from an image is a challenging problem in artificial intelligence that requires a good understanding of the visual and textual signals and the correlations between them. The…

Computation and Language · Computer Science 2020-08-07 Arushi Goel , Basura Fernando , Thanh-Son Nguyen , Hakan Bilen

Image captioning, like many tasks involving vision and language, currently relies on Transformer-based architectures for extracting the semantics in an image and translating it into linguistically coherent descriptions. Although successful,…

Computer Vision and Pattern Recognition · Computer Science 2023-08-25 Manuele Barraco , Sara Sarto , Marcella Cornia , Lorenzo Baraldi , Rita Cucchiara

Vision-language models (VLMs) achieve remarkable performance through large-scale image-text pretraining. However, their reliance on labeled image datasets limits scalability and leaves vast amounts of unlabeled image data underutilized. To…

Computer Vision and Pattern Recognition · Computer Science 2025-10-16 Sanghyun Byun , Jung Ick Guack , Mohanad Odema , Baisub Lee , Jacob Song , Woo Seong Chung
‹ Prev 1 8 9 10 Next ›