English
Related papers

Related papers: Generating Visual Stories with Grounded and Corefe…

200 papers

Human observers engage in selective information uptake when classifying visual patterns. The same is true of deep neural networks, which currently constitute the best performing artificial vision systems. Our goal is to examine the…

Computer Vision and Pattern Recognition · Computer Science 2021-08-31 Chetan Ralekar , Shubham Choudhary , Tapan Kumar Gandhi , Santanu Chaudhury

Image narrative generation is a task to create a story from an image with a subjective viewpoint. Given the importance of the subjective feelings of writers, readers, and characters in storytelling, an image narrative generation method…

Computer Vision and Pattern Recognition · Computer Science 2022-04-08 Kohei Uehara , Yusuke Mori , Yusuke Mukuta , Tatsuya Harada

Person re-identification aims to maintain the identity of an individual in diverse locations through different non-overlapping camera views. The problem is fundamentally challenging due to appearance variations resulting from differing…

Computer Vision and Pattern Recognition · Computer Science 2014-10-27 Ziming Zhang , Yuting Chen , Venkatesh Saligrama

Visual representation learning has been a cornerstone in computer vision, involving typical forms such as visual embeddings, structural symbols, and text-based representations. Despite the success of CLIP-type visual embeddings, they often…

Computer Vision and Pattern Recognition · Computer Science 2024-06-18 Yiwu Zhong , Zi-Yuan Hu , Michael R. Lyu , Liwei Wang

Robust feature representation plays significant role in visual tracking. However, it remains a challenging issue, since many factors may affect the experimental performance. The existing method which combine different features by setting…

Computer Vision and Pattern Recognition · Computer Science 2017-05-15 Yuqi Han , Chenwei Deng , Zengshuo Zhang , Jiatong Li , Baojun Zhao

Visual grounding tasks aim to localize image regions based on natural language references. In this work, we explore whether generative VLMs predominantly trained on image-text data could be leveraged to scale up the text annotation of…

Computer Vision and Pattern Recognition · Computer Science 2024-07-23 Shijie Wang , Dahun Kim , Ali Taalimi , Chen Sun , Weicheng Kuo

Generating high-quality stories spanning thousands of tokens requires competency across a variety of skills, from tracking plot and character arcs to keeping a consistent and engaging style. Due to the difficulty of sourcing labeled…

Computation and Language · Computer Science 2025-09-09 Alexander Gurung , Mirella Lapata

Storytelling and narrative are fundamental to human experience, intertwined with our social and cultural engagement. As such, researchers have long attempted to create systems that can generate stories automatically. In recent years,…

Computation and Language · Computer Science 2023-09-13 Yuxin Wang , Jieru Lin , Zhiwei Yu , Wei Hu , Börje F. Karlsson

Training-free consistent text-to-image generation depicting the same subjects across different images is a topic of widespread recent interest. Existing works in this direction predominantly rely on cross-frame self-attention; which…

Computer Vision and Pattern Recognition · Computer Science 2025-04-09 Jaskirat Singh , Junshen Kevin Chen , Jonas Kohler , Michael Cohen

Can visual artworks created using generative visual algorithms inspire human creativity in storytelling? We asked writers to write creative stories from a starting prompt, and provided them with visuals created by generative AI models from…

Human-Computer Interaction · Computer Science 2021-10-29 Safinah Ali , Devi Parikh

Despite the success of existing referenced metrics (e.g., BLEU and MoverScore), they correlate poorly with human judgments for open-ended text generation including story or dialog generation because of the notorious one-to-many issue: there…

Computation and Language · Computer Science 2020-09-17 Jian Guan , Minlie Huang

Visual grounding (VG) aims to locate a specific target in an image based on a given language query. The discriminative information from context is important for distinguishing the target from other objects, particularly for the targets that…

Computer Vision and Pattern Recognition · Computer Science 2023-12-20 Wei Tang , Liang Li , Xuejing Liu , Lu Jin , Jinhui Tang , Zechao Li

GPT-Vision has impressed us on a range of vision-language tasks, but it comes with the familiar new challenge: we have little idea of its capabilities and limitations. In our study, we formalize a process that many have instinctively been…

Computation and Language · Computer Science 2023-11-06 Alyssa Hwang , Andrew Head , Chris Callison-Burch

We present a visually-grounded language understanding model based on a study of how people verbally describe objects in scenes. The emphasis of the model is on the combination of individual word meanings to produce meanings for complex…

Artificial Intelligence · Computer Science 2011-07-04 P. Gorniak , D. Roy

Visual dialog entails answering a series of questions grounded in an image, using dialog history as context. In addition to the challenges found in visual question answering (VQA), which can be seen as one-round dialog, visual dialog…

Computer Vision and Pattern Recognition · Computer Science 2018-09-07 Satwik Kottur , José M. F. Moura , Devi Parikh , Dhruv Batra , Marcus Rohrbach

Sequential identity consistency under precise transient attribute control remains a long-standing challenge in controllable visual storytelling. Existing datasets lack sufficient fidelity and fail to disentangle stable identities from…

Computer Vision and Pattern Recognition · Computer Science 2025-12-08 Xingxi Yin , Yicheng Li , Gong Yan , Chenglin Li , Jian Zhao , Cong Huang , Yue Deng , Yin Zhang

Distributional semantic models capture word-level meaning that is useful in many natural language processing tasks and have even been shown to capture cognitive aspects of word meaning. The majority of these models are purely text based,…

Computation and Language · Computer Science 2022-03-31 Danny Merkx , Stefan L. Frank , Mirjam Ernestus

Text-to-image generation models have made significant progress in producing high-quality images from textual descriptions, yet they continue to struggle with maintaining subject consistency across multiple images, a fundamental requirement…

Computer Vision and Pattern Recognition · Computer Science 2025-11-11 Mingxiao Li , Mang Ning , Marie-Francine Moens

With the maturing of deep learning systems, trustworthiness is becoming increasingly important for model assessment. We understand trustworthiness as the combination of explainability and robustness. Generative classifiers (GCs) are a…

Computer Vision and Pattern Recognition · Computer Science 2020-12-03 Radek Mackowiak , Lynton Ardizzone , Ullrich Köthe , Carsten Rother

We introduce Prototype Generation, a stricter and more robust form of feature visualisation for model-agnostic, data-independent interpretability of image classification models. We demonstrate its ability to generate inputs that result in…

Computer Vision and Pattern Recognition · Computer Science 2023-10-02 Arush Tagade , Jessica Rumbelow
‹ Prev 1 8 9 10 Next ›