English
Related papers

Related papers: TalkMosaic: Interactive PhotoMosaic with Multi-mod…

200 papers

A creative image-and-text generative AI system mimics humans' extraordinary abilities to provide users with diverse and comprehensive caption suggestions, as well as rich image creations. In this work, we demonstrate such an AI creation…

Computer Vision and Pattern Recognition · Computer Science 2021-10-20 Yupan Huang , Bei Liu , Jianlong Fu , Yutong Lu

Humans can easily understand a single image as depicting multiple potential objects permitting interaction. We use this skill to plan our interactions with the world and accelerate understanding new objects without engaging in interaction.…

Computer Vision and Pattern Recognition · Computer Science 2023-08-09 Shengyi Qian , David F. Fouhey

Human-AI interactivity is a critical aspect that reflects the usability of multimodal large language models (MLLMs). However, existing end-to-end MLLMs only allow users to interact with them through language instructions, leading to the…

Computation and Language · Computer Science 2023-07-19 Liang Zhao , En Yu , Zheng Ge , Jinrong Yang , Haoran Wei , Hongyu Zhou , Jianjian Sun , Yuang Peng , Runpei Dong , Chunrui Han , Xiangyu Zhang

Autonomous agents operating in public spaces must consider how their behaviors might affect the humans around them, even when not directly interacting with them. To this end, it is often beneficial to be predictable and appear naturalistic.…

Multiagent Systems · Computer Science 2025-05-06 Hamzah I. Khan , David Fridovich-Keil

Visual understanding of geometric structures with complex spatial relationships is a fundamental component of human intelligence. As children, we learn how to reason about structure not only from observation, but also by interacting with…

Computer Vision and Pattern Recognition · Computer Science 2022-07-29 Aaron Walsman , Muru Zhang , Klemen Kotar , Karthik Desingh , Ali Farhadi , Dieter Fox

Human-Machine Interaction (HMI) systems have gained huge interest in recent years, with reference expression comprehension being one of the main challenges. Traditionally human-machine interaction has been mostly limited to speech and…

Human-Computer Interaction · Computer Science 2023-06-21 Aman Jain , Anirudh Reddy Kondapally , Kentaro Yamada , Hitomi Yanaka

Large Multimodal Models (LMMs) have shown strong potential for assisting users in tasks, such as programming, content creation, and information access, yet their interaction remains largely limited to traditional interfaces such as desktops…

Human-Computer Interaction · Computer Science 2026-02-12 Liuchuan Yu , Yongqi Zhang , Lap-Fai Yu

One of the key challenges faced by autistic children is understanding social affordances in complex environments, which further impacts their ability to respond appropriately to social signals. In traffic scenarios, this impairment can even…

Human-Computer Interaction · Computer Science 2025-02-06 Yancheng Cao , Yangyang HE , Yonglin Chen , Menghan Chen , Shanhe You , Yulin Qiu , Min Liu , Chuan Luo , Chen Zheng , Xin Tong , Jing Liang , Jiangtao Gong

Multimodal Vision-Language Models (VLMs) enable powerful applications from their fused understanding of images and language, but many perform poorly on UI tasks due to the lack of UI training data. In this paper, we adapt a recipe for…

Human-Computer Interaction · Computer Science 2023-10-10 Yue Jiang , Eldon Schoop , Amanda Swearngin , Jeffrey Nichols

Recently, diffusion models have achieved great success in image synthesis. However, when it comes to the layout-to-image generation where an image often has a complex scene of multiple objects, how to make strong control over both the…

Computer Vision and Pattern Recognition · Computer Science 2024-03-13 Guangcong Zheng , Xianpan Zhou , Xuewei Li , Zhongang Qi , Ying Shan , Xi Li

Service and assistive robots are increasingly being deployed in dynamic social environments; however, ensuring transparent and explainable interactions remains a significant challenge. This paper presents a multimodal explainability module…

Robotics · Computer Science 2026-04-09 Oluwadamilola Sotomi , Devika Kodi , Aliasghar Arab

Recent advancements in large language models (LLMs) have demonstrated extraordinary comprehension capabilities with remarkable breakthroughs on various vision-language tasks. However, the application of LLMs in generating reliable medical…

Artificial Intelligence · Computer Science 2025-02-18 Xueshen Li , Xinlong Hou , Ziyi Huang , Yu Gan

In this paper, we propose a novel approach to establish a connection between linguistic objects and classes in Large Language Model Machines (LLMMs) such as GPT3.5 and GPT4, and their counterparts in high level programming languages like…

Human-Computer Interaction · Computer Science 2023-04-13 Yoichi Ochiai , Naruya Kondo , Tatsuki Fushimi

The performance of ChatGPT\copyright{} and other LLMs has improved tremendously, and in online environments, they are increasingly likely to be used in a wide variety of situations, such as ChatBot on web pages, call center operations using…

Human-Computer Interaction · Computer Science 2025-02-19 Hiroki Tanioka , Tetsushi Ueta , Masahiko Sano

We propose a visual-linguistic representation learning approach within a self-supervised learning framework by introducing a new operation, loss, and data augmentation strategy. First, we generate diverse features for the image-text…

Computer Vision and Pattern Recognition · Computer Science 2023-04-04 Jaeyoo Park , Bohyung Han

Instruction tuning unlocks the superior capability of Large Language Models (LLM) to interact with humans. Furthermore, recent instruction-following datasets include images as visual inputs, collecting responses for image-based…

Computer Vision and Pattern Recognition · Computer Science 2024-02-06 Yanzhe Zhang , Ruiyi Zhang , Jiuxiang Gu , Yufan Zhou , Nedim Lipka , Diyi Yang , Tong Sun

Situated embodied conversation requires robots to interleave real-time dialogue with active perception: deciding what to look at, when to look, and what to say under tight latency constraints. We present a simple, minimal system recipe that…

Robotics · Computer Science 2026-02-05 Dong Won Lee , Sarah Gillet , Louis-Philippe Morency , Cynthia Breazeal , Hae Won Park

The recent developments in deep learning led to the integration of natural language processing (NLP) with computer vision, resulting in powerful integrated Vision and Language Models (VLMs). Despite their remarkable capabilities, these…

Computer Vision and Pattern Recognition · Computer Science 2024-10-08 Harshit , Tolga Tasdizen

In recent years, large-scale pre-trained multimodal models (LMMs) generally emerge to integrate the vision and language modalities, achieving considerable success in multimodal tasks, such as text-image classification. The growing size of…

Computer Vision and Pattern Recognition · Computer Science 2025-07-11 Xinyao Yu , Hao Sun , Zeyu Ling , Ziwei Niu , Zhenjia Bai , Rui Qin , Yen-Wei Chen , Lanfen Lin

Natural language is expected to be a key medium for various human-machine interactions in the era of large language models. When it comes to the biochemistry field, a series of tasks around molecules (e.g., property prediction, molecule…

Computation and Language · Computer Science 2023-06-22 Zheni Zeng , Bangchen Yin , Shipeng Wang , Jiarui Liu , Cheng Yang , Haishen Yao , Xingzhi Sun , Maosong Sun , Guotong Xie , Zhiyuan Liu