中文
相关论文

相关论文: ImageInWords: Unlocking Hyper-Detailed Image Descr…

200 篇论文

Recent text-guided image editing (TIE) models have achieved remarkable progress, while many edited images still suffer from issues such as artifacts, unexpected editings, unaesthetic contents. Although some benchmarks and methods have been…

计算机视觉与模式识别 · 计算机科学 2026-03-17 Zitong Xu , Huiyu Duan , Zhongpeng Ji , Xinyun Zhang , Yutao Liu , Xiongkuo Min , Ke Gu , Jian Zhang , Shusong Xu , Jinwei Chen , Bo Li , Guangtao Zhai

Multimodal large language models (MLLMs) excel at generating highly detailed captions but often produce hallucinations. Our analysis reveals that existing hallucination detection methods struggle with detailed captions. We attribute this to…

计算机视觉与模式识别 · 计算机科学 2025-07-08 Saehyung Lee , Seunghyun Yoon , Trung Bui , Jing Shi , Sungroh Yoon

Large Vision-Language Models (LVLMs), despite their recent success, are hardly comprehensively tested for their cognitive abilities. Inspired by the prevalent use of the Cookie Theft task in human cognitive tests, we propose a novel…

人工智能 · 计算机科学 2025-02-14 Xiujie Song , Mengyue Wu , Kenny Q. Zhu , Chunhao Zhang , Yanyi Chen

Text-to-image diffusion models achieved a remarkable leap in capabilities over the last few years, enabling high-quality and diverse synthesis of images from a textual prompt. However, even the most advanced models often struggle to…

计算机视觉与模式识别 · 计算机科学 2023-10-26 Eyal Segalis , Dani Valevski , Danny Lumen , Yossi Matias , Yaniv Leviathan

Vision-language models (VLMs) are increasingly used to make visual content accessible via text-based descriptions. In current systems, however, description specificity is often conflated with their length. We argue that these two concepts…

计算与语言 · 计算机科学 2026-04-21 Rhea Kapur , Robert Hawkins , Elisa Kreiss

The milestone improvements brought about by deep representation learning and pre-training techniques have led to large performance gains across downstream NLP, IR and Vision tasks. Multimodal modeling techniques aim to leverage large…

计算机视觉与模式识别 · 计算机科学 2023-02-21 Krishna Srinivasan , Karthik Raman , Jiecao Chen , Michael Bendersky , Marc Najork

In the era of large-scale visual data, understanding collections of images is a challenging yet important task. To this end, we introduce ImageSet2Text, a novel method to automatically generate natural language descriptions of image sets.…

计算机视觉与模式识别 · 计算机科学 2025-11-14 Piera Riccio , Francesco Galati , Kajetan Schweighofer , Noa Garcia , Nuria Oliver

Image Aesthetic Assessment (IAA) is a long-standing and challenging research task. However, its subset, Human Image Aesthetic Assessment (HIAA), has been scarcely explored. To bridge this research gap, our work pioneers a holistic…

计算机视觉与模式识别 · 计算机科学 2025-05-29 Zhichao Liao , Xiaokun Liu , Wenyu Qin , Qingyu Li , Qiulin Wang , Pengfei Wan , Di Zhang , Long Zeng , Pingfa Feng

Recent progress on image captioning has made it possible to generate novel sentences describing images in natural language, but compressing an image into a single sentence can describe visual content in only coarse detail. While one new…

计算机视觉与模式识别 · 计算机科学 2017-04-11 Jonathan Krause , Justin Johnson , Ranjay Krishna , Li Fei-Fei

Several services for people with visual disabilities have emerged recently due to achievements in Assistive Technologies and Artificial Intelligence areas. Despite the growth in assistive systems availability, there is a lack of services…

计算机视觉与模式识别 · 计算机科学 2022-02-17 Daniel Louzada Fernandes , Marcos Henrique Fonseca Ribeiro , Fabio Ribeiro Cerqueira , Michel Melo Silva

This paper presents a new framework for visual bag-of-words (BOW) refinement and reduction to overcome the drawbacks associated with the visual BOW model which has been widely used for image classification. Although very influential in the…

计算机视觉与模式识别 · 计算机科学 2017-07-04 Zhiwu Lu , Liwei Wang , Ji-Rong Wen

In this study, we introduce a low cost method for generating descriptions from images containing novel objects. Generally, constructing a model, which can explain images with novel objects, is costly because of the following: (1) collecting…

计算机视觉与模式识别 · 计算机科学 2020-03-09 Mikihiro Tanaka , Tatsuya Harada

Image captioning models tend to describe images in an object-centric way, emphasising visible objects. But image descriptions can also abstract away from objects and describe the type of scene depicted. In this paper, we explore the…

计算与语言 · 计算机科学 2022-11-11 Michele Cafagna , Kees van Deemter , Albert Gatt

A new class of applications based on visual search engines are emerging, especially on smart-phones that have evolved into powerful tools for processing images and videos. The state-of-the-art algorithms for large visual content recognition…

计算机视觉与模式识别 · 计算机科学 2016-04-15 Giuseppe Amato , Fabrizio Falchi , Claudio Gennaro

Understanding the multi-dimensional attributes and intensity nuances of image-evoked emotions is pivotal for advancing machine empathy and empowering diverse human-computer interaction applications. However, existing models are still…

计算机视觉与模式识别 · 计算机科学 2026-05-26 Lancheng Gao , Ziheng Jia , Zixuan Xing , Wei Sun , Huiyu Duan , Guangtao Zhai , Xiongkuo Min

Automatically generating descriptive captions for images is a well-researched area in computer vision. However, existing evaluation approaches focus on measuring the similarity between two sentences disregarding fine-grained semantics of…

计算机视觉与模式识别 · 计算机科学 2019-08-07 Philipp Harzig , Dan Zecha , Rainer Lienhart , Carolin Kaiser , René Schallner

While recent advancements in multimodal language models have enabled image generation from expressive multi-image instructions, existing methods struggle to maintain performance under complex interleaved instructions. This limitation stems…

计算机视觉与模式识别 · 计算机科学 2026-05-13 Yabo Zhang , Kunchang Li , Dewei Zhou , Xinyu Huang , Xun Wang

Image-text matching (ITM) aims to address the fundamental challenge of aligning visual and textual modalities, which inherently differ in their representations, continuous, high-dimensional image features vs. discrete, structured text. We…

多媒体 · 计算机科学 2025-07-14 Junyu Chen , Yihua Gao , Mingyong Li

Large Language Models (LLMs) have demonstrated remarkable success in tasks like the Winograd Schema Challenge (WSC), showcasing advanced textual common-sense reasoning. However, applying this reasoning to multimodal domains, where…

计算与语言 · 计算机科学 2024-06-04 Brendan Park , Madeline Janecek , Naser Ezzati-Jivan , Yifeng Li , Ali Emami

CLIP has shown impressive results in aligning images and texts at scale. However, its ability to capture detailed visual features remains limited because CLIP matches images and texts at a global level. To address this issue, we propose…

计算机视觉与模式识别 · 计算机科学 2024-12-05 Rui Xiao , Sanghwan Kim , Mariana-Iuliana Georgescu , Zeynep Akata , Stephan Alaniz