English
Related papers

Related papers: Automatic Annotation of Structured Facts in Images

200 papers

We present a self-supervised method to improve an agent's abilities in describing arbitrary objects while actively exploring a generic environment. This is a challenging problem, as current models struggle to obtain coherent image captions…

Computer Vision and Pattern Recognition · Computer Science 2025-09-18 Tommaso Galliena , Tommaso Apicella , Stefano Rosa , Pietro Morerio , Alessio Del Bue , Lorenzo Natale

Motivated by the recent progress in generative models, we introduce a model that generates images from natural language descriptions. The proposed model iteratively draws patches on a canvas, while attending to the relevant words in the…

Machine Learning · Computer Science 2016-03-01 Elman Mansimov , Emilio Parisotto , Jimmy Lei Ba , Ruslan Salakhutdinov

This project aims to create an automated image captioning system that generates natural language descriptions for input images by integrating techniques from computer vision and natural language processing. We employ various different…

Computer Vision and Pattern Recognition · Computer Science 2024-12-17 Joshua Adrian Cahyono , Jeremy Nathan Jusuf

Fact-checking has become increasingly important due to the speed with which both information and misinformation can spread in the modern media ecosystem. Therefore, researchers have been exploring how fact-checking can be automated, using…

Computation and Language · Computer Science 2022-06-07 Zhijiang Guo , Michael Schlichtkrull , Andreas Vlachos

Image aesthetic quality assessment has been a relatively hot topic during the last decade. Most recently, comments type assessment (aesthetic captions) has been proposed to describe the general aesthetic impression of an image using text.…

Computer Vision and Pattern Recognition · Computer Science 2019-07-30 Xin Jin , Le Wu , Geng Zhao , Xiaodong Li , Xiaokun Zhang , Shiming Ge , Dongqing Zou , Bin Zhou , Xinghui Zhou

We address the task of detecting foiled image captions, i.e. identifying whether a caption contains a word that has been deliberately replaced by a semantically similar word, thus rendering it inaccurate with respect to the image being…

Computer Vision and Pattern Recognition · Computer Science 2018-05-18 Pranava Madhyastha , Josiah Wang , Lucia Specia

Stylized visual captioning aims to generate image or video descriptions with specific styles, making them more attractive and emotionally appropriate. One major challenge with this task is the lack of paired stylized captions for visual…

Multimedia · Computer Science 2023-08-01 Dingyi Yang , Hongyu Chen , Xinglin Hou , Tiezheng Ge , Yuning Jiang , Qin Jin

Localizing phrases in images is an important part of image understanding and can be useful in many applications that require mappings between textual and visual information. Existing work attempts to learn these mappings from examples of…

Computer Vision and Pattern Recognition · Computer Science 2019-08-22 Josiah Wang , Lucia Specia

Advancements in Text-to-Image synthesis over recent years have focused more on improving the quality of generated samples using datasets with descriptive prompts. However, real-world image-caption pairs present in domains such as news data…

Computer Vision and Pattern Recognition · Computer Science 2024-08-06 Aashish Anantha Ramakrishnan , Sharon X. Huang , Dongwon Lee

We explore the use of a knowledge graphs, that capture general or commonsense knowledge, to augment the information extracted from images by the state-of-the-art methods for image captioning. The results of our experiments, on several…

Computer Vision and Pattern Recognition · Computer Science 2019-01-28 Yimin Zhou , Yiwei Sun , Vasant Honavar

Visual question answering is concerned with answering free-form questions about an image. Since it requires a deep linguistic understanding of the question and the ability to associate it with various objects that are present in the image,…

Machine Learning · Computer Science 2020-07-03 Marcel Hildebrandt , Hang Li , Rajat Koner , Volker Tresp , Stephan Günnemann

Gaze reflects how humans process visual scenes and is therefore increasingly used in computer vision systems. Previous works demonstrated the potential of gaze for object-centric tasks, such as object localization and recognition, but it…

Computer Vision and Pattern Recognition · Computer Science 2016-08-19 Yusuke Sugano , Andreas Bulling

We present a framework for automatically reconfiguring images of street scenes by populating, depopulating, or repopulating them with objects such as pedestrians or vehicles. Applications of this method include anonymizing images to enhance…

Computer Vision and Pattern Recognition · Computer Science 2021-03-31 Yifan Wang , Andrew Liu , Richard Tucker , Jiajun Wu , Brian L. Curless , Steven M. Seitz , Noah Snavely

Significant progress has been made in spatial intelligence, spanning both spatial reconstruction and world exploration. However, the scalability and real-world fidelity of current models remain severely constrained by the scarcity of…

Computer Vision and Pattern Recognition · Computer Science 2025-12-19 Jiahao Wang , Yufeng Yuan , Rujie Zheng , Youtian Lin , Jian Gao , Lin-Zhuo Chen , Yajie Bao , Yi Zhang , Chang Zeng , Yanxi Zhou , Xiao-Xiao Long , Hao Zhu , Zhaoxiang Zhang , Xun Cao , Yao Yao

With the emergence of collaborative robots (cobots), human-robot collaboration in industrial manufacturing is coming into focus. For a cobot to act autonomously and as an assistant, it must understand human actions during assembly. To…

Robotics · Computer Science 2023-04-18 Dustin Aganian , Benedict Stephan , Markus Eisenbach , Corinna Stretz , Horst-Michael Gross

We present a method for estimating articulated human pose from a single static image based on a graphical model with novel pairwise relations that make adaptive use of local image measurements. More precisely, we specify a graphical model…

Computer Vision and Pattern Recognition · Computer Science 2014-11-05 Xianjie Chen , Alan Yuille

In this paper we propose the construction of linguistic descriptions of images. This is achieved through the extraction of scene description graphs (SDGs) from visual scenes using an automatically constructed knowledge base. SDGs are…

Computer Vision and Pattern Recognition · Computer Science 2015-11-12 Somak Aditya , Yezhou Yang , Chitta Baral , Cornelia Fermuller , Yiannis Aloimonos

This paper introduces a novel approach to enhance existing motion captioning methods, which directly map representations of movement to high-level descriptive captions (e.g., ``a person doing jumping jacks"). The existing methods require…

Machine Learning · Computer Science 2025-09-03 Clayton Leite , Yu Xiao

Natural language plays a critical role in many computer vision applications, such as image captioning, visual question answering, and cross-modal retrieval, to provide fine-grained semantic information. Unfortunately, while human pose is…

Computer Vision and Pattern Recognition · Computer Science 2024-09-11 Ginger Delmas , Philippe Weinzaepfel , Thomas Lucas , Francesc Moreno-Noguer , Grégory Rogez

Textual scene graph parsing has become increasingly important in various vision-language applications, including image caption evaluation and image retrieval. However, existing scene graph parsers that convert image captions into scene…

Computation and Language · Computer Science 2023-06-02 Zhuang Li , Yuyang Chai , Terry Yue Zhuo , Lizhen Qu , Gholamreza Haffari , Fei Li , Donghong Ji , Quan Hung Tran