中文
相关论文

相关论文: Not (yet) the whole story: Evaluating Visual Story…

200 篇论文

Quantitative evaluation metrics have traditionally been pivotal in gauging the advancements of artificial intelligence systems, including large language models (LLMs). However, these metrics have inherent limitations. Given the intricate…

Digital video is central to communication, education, and entertainment, but without audio description (AD), blind and low-vision users are excluded. While crowdsourced platforms and vision-language models (VLMs) expand AD production,…

Text-to-image generation models have made significant progress in producing high-quality images from textual descriptions, yet they continue to struggle with maintaining subject consistency across multiple images, a fundamental requirement…

计算机视觉与模式识别 · 计算机科学 2025-11-11 Mingxiao Li , Mang Ning , Marie-Francine Moens

Natural language descriptions sometimes accompany visualizations to better communicate and contextualize their insights, and to improve their accessibility for readers with disabilities. However, it is difficult to evaluate the usefulness…

人机交互 · 计算机科学 2021-10-12 Alan Lundgard , Arvind Satyanarayan

Vision Language Models (VLMs) are impressive at visual question answering and image captioning. But they underperform on multi-step visual reasoning -- even compared to LLMs on the same tasks presented in text form -- giving rise to…

计算机视觉与模式识别 · 计算机科学 2025-06-03 Simon Park , Abhishek Panigrahi , Yun Cheng , Dingli Yu , Anirudh Goyal , Sanjeev Arora

Bridging vision and natural language is a longstanding goal in computer vision and multimedia research. While earlier works focus on generating a single-sentence description for visual content, recent works have studied paragraph…

多媒体 · 计算机科学 2020-05-15 Junnan Li , Yongkang Wong , Qi Zhao , Mohan S. Kankanhalli

Language models (LM) are very powerful in lipreading systems. Language models built upon the ground truth utterances of datasets learn grammar and structure rules of words and sentences (the latter in the case of continuous speech).…

音频与语音处理 · 电气工程与系统科学 2018-09-19 Helen L Bear

A major challenge in visually grounded language generation is to build robust benchmark datasets and models that can generalize well in real-world settings. To do this, it is critical to ensure that our evaluation protocols are correct, and…

计算与语言 · 计算机科学 2020-10-09 Wanrong Zhu , Xin Eric Wang , Pradyumna Narayana , Kazoo Sone , Sugato Basu , William Yang Wang

Characters are essential to the plot of any story. Establishing the characters before writing a story can improve the clarity of the plot and the overall flow of the narrative. However, previous work on visual storytelling tends to focus on…

计算与语言 · 计算机科学 2023-04-03 Danyang Liu , Frank Keller

Vision-Language Models (VLMs) building upon the foundation of powerful large language models have made rapid progress in reasoning across visual and textual data. While VLMs perform well on vision tasks that they are trained on, our results…

计算机视觉与模式识别 · 计算机科学 2025-06-26 Zixuan Wu , Yoolim Kim , Carolyn Jane Anderson

Translating natural language to visualization (NL2VIS) has shown great promise for visual data analysis, but it remains a challenging task that requires multiple low-level implementations, such as natural language processing and…

人机交互 · 计算机科学 2024-08-08 Nan Chen , Yuge Zhang , Jiahang Xu , Kan Ren , Yuqing Yang

Effective planning requires strong world models, but high-level world models that can understand and reason about actions with semantic and temporal abstraction remain largely underdeveloped. We introduce the Vision Language World Model…

人工智能 · 计算机科学 2025-09-09 Delong Chen , Theo Moutakanni , Willy Chung , Yejin Bang , Ziwei Ji , Allen Bolourchi , Pascale Fung

Automated storytelling has long captured the attention of researchers for the ubiquity of narratives in everyday life. However, it is challenging to maintain coherence and stay on-topic toward a specific ending when generating narratives…

计算与语言 · 计算机科学 2022-05-17 Xiangyu Peng , Kaige Xie , Amal Alabdulkarim , Harshith Kayam , Samihan Dani , Mark O. Riedl

Large language models (LLMs) can be employed for automating the generation of software requirements from natural language inputs such as the transcripts of elicitation interviews. However, evaluating whether those derived requirements…

计算与语言 · 计算机科学 2025-10-13 Francesco Dente , Fabiano Dalpiaz , Paolo Papotti

With rapid advances in large language models (LLMs), there has been an increasing application of LLMs in creative content ideation and generation. A critical question emerges: can current LLMs provide ideas that are diverse enough to truly…

计算与语言 · 计算机科学 2025-09-03 Weijia Xu , Nebojsa Jojic , Sudha Rao , Chris Brockett , Bill Dolan

Beyond conventional paradigms of translating speech and text, recently, there has been interest in automated transcreation of images to facilitate localization of visual content across different cultures. Attempts to define this as a formal…

计算与语言 · 计算机科学 2025-03-24 Simran Khanuja , Vivek Iyer , Claire He , Graham Neubig

A fundamental characteristic common to both human vision and natural language is their compositional nature. Yet, despite the performance gains contributed by large vision and language pretraining, we find that: across 7 architectures…

计算与语言 · 计算机科学 2023-05-17 Zixian Ma , Jerry Hong , Mustafa Omer Gul , Mona Gandhi , Irena Gao , Ranjay Krishna

Vision-language models (VLMs) excel at extracting and reasoning about information from images. Yet, their capacity to leverage internal knowledge about specific entities remains underexplored. This work investigates the disparity in model…

计算与语言 · 计算机科学 2026-01-06 Ido Cohen , Daniela Gottesman , Mor Geva , Raja Giryes

What information is sufficient to learn the full richness of human scene understanding? The distributional hypothesis holds that the statistical co-occurrence of language and images captures the conceptual knowledge underlying visual…

计算机视觉与模式识别 · 计算机科学 2026-03-30 Gillian Rosenberg , Skylar Stadhard , Bruce C. Hansen , Michelle R. Greene

Traditional automated metrics for evaluating conditional natural language generation use pairwise comparisons between a single generated text and the best-matching gold-standard ground truth text. When multiple ground truths are available,…

计算与语言 · 计算机科学 2022-09-30 David M Chan , Yiming Ni , David A Ross , Sudheendra Vijayanarasimhan , Austin Myers , John Canny