中文
相关论文

相关论文: VisImages: A Fine-Grained Expert-Annotated Visuali…

200 篇论文

The arXiv has collected 1.5 million pre-print articles over 28 years, hosting literature from scientific fields including Physics, Mathematics, and Computer Science. Each pre-print features text, figures, authors, citations, categories, and…

信息检索 · 计算机科学 2019-05-02 Colin B. Clement , Matthew Bierbaum , Kevin P. O'Keeffe , Alexander A. Alemi

Image captioning is a multimodal problem that has drawn extensive attention in both the natural language processing and computer vision community. In this paper, we present a novel image captioning architecture to better explore semantics…

计算机视觉与模式识别 · 计算机科学 2020-06-23 Zhan Shi , Xu Zhou , Xipeng Qiu , Xiaodan Zhu

Current captioning datasets focus on object-centric captions, describing the visible objects in the image, e.g. "people eating food in a park". Although these datasets are useful to evaluate the ability of Vision & Language models to…

计算与语言 · 计算机科学 2023-09-26 Michele Cafagna , Kees van Deemter , Albert Gatt

Multiple benchmarks have been developed to assess the alignment between deep neural networks (DNNs) and human vision. In almost all cases these benchmarks are observational in the sense they are composed of behavioural and brain responses…

Recognizing the layout of unstructured digital documents is an important step when parsing the documents into structured machine-readable format for downstream applications. Deep neural networks that are developed for computer vision have…

计算与语言 · 计算机科学 2019-08-22 Xu Zhong , Jianbin Tang , Antonio Jimeno Yepes

We present Visual Lexicon, a novel visual language that encodes rich image information into the text space of vocabulary tokens while retaining intricate visual details that are often challenging to convey in natural language. Unlike…

计算机视觉与模式识别 · 计算机科学 2024-12-10 XuDong Wang , Xingyi Zhou , Alireza Fathi , Trevor Darrell , Cordelia Schmid

Given the central role of charts in scientific, business, and communication contexts, enhancing the chart understanding capabilities of vision-language models (VLMs) has become increasingly critical. A key limitation of existing VLMs lies…

计算机视觉与模式识别 · 计算机科学 2025-10-17 Jiangning Zhu , Yuxing Zhou , Zheng Wang , Juntao Yao , Yima Gu , Yuhui Yuan , Shixia Liu

The rapid influx of low-quality data visualisations is one of the main challenges in today's communication. Misleading, unreadable, or confusing visualisations spread misinformation, failing to fulfill their purpose. The lack of proper…

人机交互 · 计算机科学 2023-04-17 Jan Sawicki , Michał Burdukiewicz

Data visualization should be accessible for all analysts with data, not just the few with technical expertise. Visualization recommender systems aim to lower the barrier to exploring basic visualizations by automatically generating results…

人机交互 · 计算机科学 2018-08-16 Kevin Z. Hu , Michiel A. Bakker , Stephen Li , Tim Kraska , César A. Hidalgo

Transforming recorded videos into concise and accurate textual summaries is a growing challenge in multimodal learning. This paper introduces VISTA, a dataset specifically designed for video-to-text summarization in scientific domains.…

计算与语言 · 计算机科学 2025-05-27 Dongqi Liu , Chenxi Whitehouse , Xi Yu , Louis Mahon , Rohit Saxena , Zheng Zhao , Yifu Qiu , Mirella Lapata , Vera Demberg

With the introduction of the Visualization for Communication workshop (VisComm) at IEEE VIS and in light of the COVID-19 pandemic, there has been renewed interest in studying visualization as a medium of communication. However the…

人机交互 · 计算机科学 2025-05-21 Vedanshi Chetan Shah , Ab Mosca

Image captioning implies automatically generating textual descriptions of images based only on the visual input. Although this has been an extensively addressed research topic in recent years, not many contributions have been made in the…

计算机视觉与模式识别 · 计算机科学 2021-02-09 Eva Cetinic

Textbooks are one of the main mediums for delivering high-quality education to students. In particular, explanatory and illustrative visuals play a key role in retention, comprehension and general transfer of knowledge. However, many…

计算机视觉与模式识别 · 计算机科学 2023-10-23 Janvijay Singh , Vilém Zouhar , Mrinmaya Sachan

Fashion image generation has so far focused on narrow tasks such as virtual try-on, where garments appear in clean studio environments. In contrast, editorial fashion presents garments through dynamic poses, diverse locations, and carefully…

计算机视觉与模式识别 · 计算机科学 2025-10-02 Yannick Hauri , Luca A. Lanzendörfer , Till Aczel

General visualization recommendation systems typically make design decisions for the dataset automatically. However, most of them can only prune meaningless visualizations but fail to recommend targeted results. This paper contributes…

人机交互 · 计算机科学 2022-09-15 Leixian Shen , Enya Shen , Zhiwei Tai , Yihao Xu , Jianmin Wang

Image captioning models have been able to generate grammatically correct and human understandable sentences. However most of the captions convey limited information as the model used is trained on datasets that do not caption all possible…

计算机视觉与模式识别 · 计算机科学 2020-03-30 Pranav Agarwal , Alejandro Betancourt , Vana Panagiotou , Natalia Díaz-Rodríguez

While there has been remarkable progress in the performance of visual recognition algorithms, the state-of-the-art models tend to be exceptionally data-hungry. Large labeled training datasets, expensive and tedious to produce, are required…

计算机视觉与模式识别 · 计算机科学 2016-06-07 Fisher Yu , Ari Seff , Yinda Zhang , Shuran Song , Thomas Funkhouser , Jianxiong Xiao

This paper is a call to action for research and discussion on data visualization education. As visualization evolves and spreads through our professional and personal lives, we need to understand how to support and empower a broad and…

The impact of culture in visual emotion perception has recently captured the attention of multimedia research. In this study, we pro- vide powerful computational linguistics tools to explore, retrieve and browse a dataset of 16K…

计算与语言 · 计算机科学 2016-06-09 Nikolaos Pappas , Miriam Redi , Mercan Topkara , Brendan Jou , Hongyi Liu , Tao Chen , Shih-Fu Chang

Human head detection, keypoint estimation, and 3D head model fitting are essential tasks with many applications. However, traditional real-world datasets often suffer from bias, privacy, and ethical concerns, and they have been recorded in…

计算机视觉与模式识别 · 计算机科学 2024-12-06 Orest Kupyn , Eugene Khvedchenia , Christian Rupprecht