中文
相关论文

相关论文: Testing Relational Understanding in Text-Guided Im…

200 篇论文

Deep learning has been shown to achieve impressive results in several tasks where a large amount of training data is available. However, deep learning solely focuses on the accuracy of the predictions, neglecting the reasoning process…

人工智能 · 计算机科学 2020-02-07 Giuseppe Marra , Michelangelo Diligenti , Francesco Giannini , Marco Gori , Marco Maggini

Large text-to-image models have achieved astonishing performance in synthesizing diverse and high-quality images guided by texts. With detail-oriented conditioning control, even finer-grained spatial control can be achieved. However, some…

计算机视觉与模式识别 · 计算机科学 2024-03-19 Yuhe Liu , Mengxue Kang , Zengchang Qin , Xiangxiang Chu

We propose an end-to-end network for image generation from given structured-text that consists of the visual-relation layout module and the pyramid of GANs, namely stacking-GANs. Our visual-relation layout module uses relations among…

计算机视觉与模式识别 · 计算机科学 2020-07-21 Duc Minh Vo , Akihiro Sugimoto

Reasoning about the relationships between object pairs in images is a crucial task for holistic scene understanding. Most of the existing works treat this task as a pure visual classification task: each type of relationship or phrase is…

计算机视觉与模式识别 · 计算机科学 2017-11-22 Wentong Liao , Lin Shuai , Bodo Rosenhahn , Michael Ying Yang

Deep generative models have shown impressive results in text-to-image synthesis. However, current text-to-image models often generate images that are inadequately aligned with text prompts. We propose a fine-tuning method for aligning such…

The human cognitive system exhibits remarkable flexibility and generalization capabilities, partly due to its ability to form low-dimensional, compositional representations of the environment. In contrast, standard neural network…

人工智能 · 计算机科学 2024-02-29 Declan Campbell , Jonathan D. Cohen

It is important for machines to interpret human emotions properly for better human-machine communications, as emotion is an essential part of human-to-human communications. One aspect of emotion is reflected in the language we use. How to…

计算与语言 · 计算机科学 2018-08-23 Ji Ho Park

Visual relation detection (VRD) aims to identify relationships (or interactions) between object pairs in an image. Although recent VRD models have achieved impressive performance, they are all restricted to pre-defined relation categories,…

计算机视觉与模式识别 · 计算机科学 2025-04-17 Kaifeng Gao , Siqi Chen , Hanwang Zhang , Jun Xiao , Yueting Zhuang , Qianru Sun

As the intermediate-level representations bridging the two levels, structured representations of visual scenes, such as visual relationships between pairwise objects, have been shown to not only benefit compositional models in learning to…

计算机视觉与模式识别 · 计算机科学 2022-07-12 Meng-Jiun Chiou

Visual communication, dating back to prehistoric cave paintings, is the use of visual elements to convey ideas and information. In today's visually saturated world, effective design demands an understanding of graphic design principles,…

计算机视觉与模式识别 · 计算机科学 2024-12-02 Yael Vinker

Logic reasoning is a significant ability of human intelligence and also an important task in artificial intelligence. The existing logic reasoning methods, quite often, need to design some reasoning patterns beforehand. This has led to an…

计算机视觉与模式识别 · 计算机科学 2021-06-30 Qian Guo , Yuhua Qian , Xinyan Liang , Yanhong She , Deyu Li , Jiye Liang

This paper studies the task of full generative modelling of realistic images of humans, guided only by coarse sketch of the pose, while providing control over the specific instance or type of outfit worn by the user. This is a difficult…

计算机视觉与模式识别 · 计算机科学 2019-06-06 Xu Chen , Jie Song , Otmar Hilliges

Unified multimodal understanding and generation models recently have achieve significant improvement in image generation capability, yet a large gap remains in instruction following and detail preservation compared to systems that tightly…

Language models (LMs) are trained on collections of documents, written by individual human agents to achieve specific goals in an outside world. During training, LMs have access only to text of these documents, with no direct evidence of…

计算与语言 · 计算机科学 2022-12-06 Jacob Andreas

Image and shape editing are ubiquitous among digital artworks. Graphics algorithms facilitate artists and designers to achieve desired editing intents without going through manually tedious retouching. In the recent advance of machine…

图形学 · 计算机科学 2023-04-20 Cheng-Kang Ted Chao , Yotam Gingold

Accurately generating images of human bodies from text remains a challenging problem for state of the art text-to-image models. Commonly observed body-related artifacts include extra or missing limbs, unrealistic poses, blurred body parts,…

计算机视觉与模式识别 · 计算机科学 2024-12-09 Nefeli Andreou , Varsha Vivek , Ying Wang , Alex Vorobiov , Tiffany Deng , Raja Bala , Larry Davis , Betty Mohler Tesch

Current word embedding models despite their success, still suffer from their lack of grounding in the real world. In this line of research, Gunther et al. 2022 proposed a behavioral experiment to investigate the relationship between words…

计算与语言 · 计算机科学 2023-11-01 Hassan Shahmohammadi , Maria Heitmeier , Elnaz Shafaei-Bajestan , Hendrik P. A. Lensch , Harald Baayen

Human-AI collaboration requires AI agents to understand human behavior for effective coordination. While advances in foundation models show promising capabilities in understanding and showing human-like behavior, their application in…

机器人学 · 计算机科学 2026-05-07 Shinas Shaji , Teena Chakkalayil Hassan , Sebastian Houben , Alex Mitrevski

The perceptual representations supporting our ability to recognize faces remain a computational mystery. Deep neural networks offer mechanistic hypotheses for human face perception, but theoretically distinct models often make…

神经元与认知 · 定量生物学 2026-05-14 Wenxuan Guo , Heiko H. Schütt , Kamila Maria Jozwik , Katherine R. Storrs , Nikolaus Kriegeskorte , Tal Golan

World models have garnered substantial interest in the AI community. These are internal representations that simulate aspects of the external world, track entities and states, capture causal relationships, and enable prediction of…

人工智能 · 计算机科学 2025-11-18 Tarun Gupta , Danish Pruthi