中文
相关论文

相关论文: JamendoMaxCaps: A Large Scale Music-caption Datase…

200 篇论文

While advanced image captioning systems are increasingly describing images coherently and exactly, recent progress in continual learning allows deep learning models to avoid catastrophic forgetting. However, the domain where image…

计算机视觉与模式识别 · 计算机科学 2020-04-22 Giang Nguyen , Tae Joon Jun , Trung Tran , Tolcha Yalew , Daeyoung Kim

Interest in the research areas related to meme propagation and generation has been increasing rapidly in the last couple of years. Meme datasets available online are either specific to a context or contain no class information. Here, we…

计算与语言 · 计算机科学 2019-12-04 Suryatej Reddy Vyalla , Vishaal Udandarao , Tanmoy Chakraborty

High-quality image captions play a crucial role in improving the performance of cross-modal applications such as text-to-image generation, text-to-video generation, and text-image retrieval. To generate long-form, high-quality captions,…

计算机视觉与模式识别 · 计算机科学 2025-04-10 Ruotian Peng , Haiying He , Yake Wei , Yandong Wen , Di Hu

Medical image captioning via vision-language models has shown promising potential for clinical diagnosis assistance. However, generating contextually relevant descriptions with accurate modality recognition remains challenging. We present…

计算机视觉与模式识别 · 计算机科学 2025-04-15 Yining Zhao , Ali Braytee , Mukesh Prasad

Visual captioning benchmarks have become outdated with the emergence of modern multimodal large language models (MLLMs), as the brief ground-truth sentences and traditional metrics fail to assess detailed captions effectively. While recent…

计算机视觉与模式识别 · 计算机科学 2025-11-27 Zhihang Liu , Chen-Wei Xie , Bin Wen , Feiwu Yu , Jixuan Chen , Pandeng Li , Boqiang Zhang , Nianzu Yang , Yinglu Li , Zuan Gao , Yun Zheng , Hongtao Xie

Image Difference Captioning (IDC) generates natural language descriptions that precisely identify differences between two images, serving as a key benchmark for fine-grained change perception, cross-modal reasoning, and image editing data…

计算机视觉与模式识别 · 计算机科学 2026-05-07 Yuancheng Wei , Haojie Zhang , Linli Yao , Lei Li , Jiali Chen , Tao Huang , Yiting Lu , Duojun Huang , Xin Li , Zhao Zhong

In this paper, we address a fundamental gap between pre-training and fine-tuning of deep neural networks: while pre-training has shifted from unimodal to multimodal learning with enhanced visual understanding, fine-tuning predominantly…

计算机视觉与模式识别 · 计算机科学 2026-01-30 Shohei Enomoto , Shin'ya Yamaguchi

Humor, deeply rooted in societal meanings and cultural details, poses a unique challenge for machines. While advances have been made in natural language processing, real-world humor often thrives in a multi-modal context, encapsulated…

计算机视觉与模式识别 · 计算机科学 2025-06-12 Yuyan Chen , Songzhou Yan , Zhihong Zhu , Zhixu Li , Yanghua Xiao

This paper introduces VideoMind, a video-centric omni-modal dataset designed for deep video content cognition and enhanced multi-modal feature representation. The dataset comprises 103K video samples (3K reserved for testing), each paired…

计算机视觉与模式识别 · 计算机科学 2025-07-25 Baoyao Yang , Wanyun Li , Dixin Chen , Junxiang Chen , Wenbin Yao , Haifeng Lin

We report on CMU Informedia Lab's system used in Google's YouTube 8 Million Video Understanding Challenge. In this multi-label video classification task, our pipeline achieved 84.675% and 84.662% GAP on our evaluation split and the official…

计算机视觉与模式识别 · 计算机科学 2017-07-26 Po-Yao Huang , Ye Yuan , Zhenzhong Lan , Lu Jiang , Alexander G. Hauptmann

Since the SciCap datasets launch in 2021, the research community has made significant progress in generating captions for scientific figures in scholarly articles. In 2023, the first SciCap Challenge took place, inviting global teams to use…

Cover song detection is a very relevant task in Music Information Retrieval (MIR) studies and has been mainly addressed using audio-based systems. Despite its potential impact in industrial contexts, low performances and lack of scalability…

信息检索 · 计算机科学 2018-08-31 Albin Andrew Correya , Romain Hennequin , Mickaël Arcos

We introduce CLaMP: Contrastive Language-Music Pre-training, which learns cross-modal representations between natural language and symbolic music using a music encoder and a text encoder trained jointly with a contrastive loss. To pre-train…

声音 · 计算机科学 2023-10-19 Shangda Wu , Dingyao Yu , Xu Tan , Maosong Sun

Despite significant advancements in multi-label text classification, the ability of existing models to generalize to novel and seldom-encountered complex concepts, which are compositions of elementary ones, remains underexplored. This…

计算与语言 · 计算机科学 2023-12-21 Yuyang Chai , Zhuang Li , Jiahui Liu , Lei Chen , Fei Li , Donghong Ji , Chong Teng

This work present a music dataset named MusicTM-Dataset, which is utilized in improving the representation learning ability of different types of cross-modal retrieval (CMR). Little large music dataset including three modalities is…

声音 · 计算机科学 2021-05-10 Donghuo Zeng , Yi Yu , Keizo Oyama

The quality of the data and annotation upper-bounds the quality of a downstream model. While there exist large text corpora and image-text pairs, high-quality video-text data is much harder to collect. First of all, manual labeling is more…

Generating detailed captions comprehending text-rich visual content in images has received growing attention for Large Vision-Language Models (LVLMs). However, few studies have developed benchmarks specifically tailored for detailed…

计算机视觉与模式识别 · 计算机科学 2025-03-27 Fan Lu , Wei Wu , Kecheng Zheng , Shuailei Ma , Biao Gong , Jiawei Liu , Wei Zhai , Yang Cao , Yujun Shen , Zheng-Jun Zha

Multimodal Large Language Models (MLLMs) have achieved remarkable visual reasoning abilities in natural images, text-rich documents, and graphic designs. However, their ability to interpret music sheets remains underexplored. To bridge this…

计算机视觉与模式识别 · 计算机科学 2025-08-15 Jian Chen , Wenye Ma , Penghang Liu , Wei Wang , Tengwei Song , Ming Li , Chenguang Wang , Jiayu Qin , Ruiyi Zhang , Changyou Chen

Video captions play a crucial role in text-to-video generation tasks, as their quality directly influences the semantic coherence and visual fidelity of the generated videos. Although large vision-language models (VLMs) have demonstrated…

计算机视觉与模式识别 · 计算机科学 2025-05-30 Shi-Xue Zhang , Hongfa Wang , Duojun Huang , Xin Li , Xiaobin Zhu , Xu-Cheng Yin

This paper explores the usage of multimodal image-to-text models to enhance text-based item retrieval. We propose utilizing pre-trained image captioning and tagging models, such as instructBLIP and CLIP, to generate text-based product…

信息检索 · 计算机科学 2024-02-14 Jason Tang , Garrin McGoldrick , Marie Al-Ghossein , Ching-Wei Chen