English
Related papers

Related papers: How Good is Google Bard's Visual Understanding? An…

200 papers

One of the key goals of artificial intelligence (AI) is the development of a multimodal system that facilitates communication with the visual world (image and video) using a natural language query. Earlier works on medical question…

Computer Vision and Pattern Recognition · Computer Science 2024-12-17 Deepak Gupta , Dina Demner-Fushman

Social intelligence, the ability to interpret emotions, intentions, and behaviors, is essential for effective communication and adaptive responses. As robots and AI systems become more prevalent in caregiving, healthcare, and education, the…

Computer Vision and Pattern Recognition · Computer Science 2025-03-28 Erika Mori , Yue Qiu , Hirokatsu Kataoka , Yoshimitsu Aoki

In this short paper, we present work evaluating an AI agent's understanding of spoken conversations about data visualizations in an online meeting scenario. There is growing interest in the development of AI-assistants that support…

Human-Computer Interaction · Computer Science 2025-10-07 Rizul Sharma , Tianyu Jiang , Seokki Lee , Jillian Aurisano

In the rapidly evolving landscape of human-computer interaction, the integration of vision capabilities into conversational agents stands as a crucial advancement. This paper presents an initial implementation of a dialogue manager that…

Robotics · Computer Science 2025-04-14 Giulio Antonio Abbo , Tony Belpaeme

Lately, researchers in artificial intelligence have been really interested in how language and vision come together, giving rise to the development of multimodal models that aim to seamlessly integrate textual and visual information.…

Computer Vision and Pattern Recognition · Computer Science 2024-10-29 Rajat Chawla , Arkajit Datta , Tushar Verma , Adarsh Jha , Anmol Gautam , Ayush Vatsal , Sukrit Chaterjee , Mukunda NS , Ishaan Bhola

Currently, dialogue systems have achieved high performance in processing text-based communication. However, they have not yet effectively incorporated visual information, which poses a significant challenge. Furthermore, existing models…

Computation and Language · Computer Science 2023-12-19 Viktor Moskvoretskii , Anton Frolov , Denis Kuznetsov

OpenAI's multimodal GPT-4o has demonstrated remarkable capabilities in image generation and editing, yet its ability to achieve world knowledge-informed semantic synthesis--seamlessly integrating domain knowledge, contextual reasoning, and…

Computer Vision and Pattern Recognition · Computer Science 2025-04-14 Ning Li , Jingran Zhang , Justin Cui

Charts and graphs help people analyze data, but can they also be useful to AI systems? To investigate this question, we perform a series of experiments with two commercial vision-language models: GPT 4.1 and Claude 3.5. Across three…

Artificial Intelligence · Computer Science 2025-07-25 Victoria R. Li , Johnathan Sun , Martin Wattenberg

Virtual assistants are becoming increasingly important speech-driven Information Retrieval platforms that assist users with various tasks. We discuss open problems and challenges with respect to modeling spoken information queries for…

Information Retrieval · Computer Science 2023-04-27 Christophe Van Gysel

Recent advancements in Natural Language Processing (NLP), particularly in Large Language Models (LLMs), associated with deep learning-based computer vision techniques, have shown substantial potential for automating a variety of tasks. One…

Computer Vision and Pattern Recognition · Computer Science 2023-07-06 Lucas Prado Osco , Eduardo Lopes de Lemos , Wesley Nunes Gonçalves , Ana Paula Marques Ramos , José Marcato Junior

Vision language models (VLMs) have advanced graphical user interface (GUI) task automation but still lag behind humans. We hypothesize this gap stems from missing core GUI knowledge, which existing training schemes (such as supervised fine…

Artificial Intelligence · Computer Science 2026-02-10 Chenrui Shi , Zedong Yu , Zhi Gao , Ruining Feng , Enqi Liu , Yuwei Wu , Yunde Jia , Liuyu Xiang , Zhaofeng He , Qing Li

Conversational generative AI systems such as ChatGPT are transforming how people seek and engage with information online. Unlike traditional search engines, these systems support open-ended, conversational inquiry, yet it remains unclear…

Human-Computer Interaction · Computer Science 2026-04-14 Yulin Yu , Yizhou Li , Siddharth Suri , Scott Counts

Multimodal large language models (MLLMs) have achieved impressive progress on vision language benchmarks, yet their capacity for visual cognitive and visuospatial reasoning remains less understood. We introduce "Mind's Eye", a…

Computer Vision and Pattern Recognition · Computer Science 2026-04-20 Rohit Sinha , Aditya Kanade , Sai Srinivas Kancheti , Vineeth N Balasubramanian , Tanuja Ganu

The Internet of Battlefield Things (IoBT) will advance the operational effectiveness of infantry units. However, this requires autonomous assets such as sensors, drones, combat equipment, and uncrewed vehicles to collaborate, securely share…

Cryptography and Security · Computer Science 2022-08-04 Sai Sree Laya Chukkapalli , Anupam Joshi , Tim Finin , Robert F. Erbacher

Decoding of seen visual contents with non-invasive brain recordings has important scientific and practical values. Efforts have been made to recover the seen images from brain signals. However, most existing approaches cannot faithfully…

Computer Vision and Pattern Recognition · Computer Science 2023-09-28 Jiaxuan Chen , Yu Qi , Yueming Wang , Gang Pan

A wide variety of NLP applications, such as machine translation, summarization, and dialog, involve text generation. One major challenge for these applications is how to evaluate whether such generated texts are actually fluent, accurate,…

Computation and Language · Computer Science 2021-10-28 Weizhe Yuan , Graham Neubig , Pengfei Liu

The Semantic Robot Vision Competition provided an excellent opportunity for our research lab to integrate our many ideas under one umbrella, inspiring both collaboration and new research. The task, visual search for an unknown object, is…

Computer Vision and Pattern Recognition · Computer Science 2009-08-20 Scott Helmer , David Meger , Pooja Viswanathan , Sancho McCann , Matthew Dockrey , Pooyan Fazli , Tristram Southey , Marius Muja , Michael Joya , Jim Little , David Lowe , Alan Mackworth

Understanding images and text together is an important aspect of cognition and building advanced Artificial Intelligence (AI) systems. As a community, we have achieved good benchmarks over language and vision domains separately, however…

Computer Vision and Pattern Recognition · Computer Science 2020-11-19 Shailaja Keyur Sampat , Yezhou Yang , Chitta Baral

Artificial Intelligence makes great advances today and starts to bridge the gap between vision and language. However, we are still far from understanding, explaining and controlling explicitly the visual content from a linguistic…

Artificial Intelligence · Computer Science 2023-09-19 Mihai Masala , Nicolae Cudlenco , Traian Rebedea , Marius Leordeanu

This study investigates the performance of eight large multimodal model (LMM)-based chatbots on the Test of Understanding Graphs in Kinematics (TUG-K), a research-based concept inventory. Graphs are a widely used representation in STEM and…

Physics Education · Physics 2024-10-25 Giulia Polverini , Bor Gregorcic