English
Related papers

Related papers: ProMQA-Assembly: Multimodal Procedural QA Dataset …

200 papers

Multimodal large language models (LLMs) have achieved notable success across various domains, while research in the medical field has largely focused on unimodal images. Meanwhile, current general-domain multimodal models for videos still…

Computer Vision and Pattern Recognition · Computer Science 2024-08-16 Jiajie Li , Garrett Skinner , Gene Yang , Brian R Quaranto , Steven D Schwaitzberg , Peter C W Kim , Jinjun Xiong

Lecture slide presentations, a sequence of pages that contain text and figures accompanied by speech, are constructed and presented carefully in order to optimally transfer knowledge to students. Previous studies in multimedia and…

Artificial Intelligence · Computer Science 2022-08-18 Dong Won Lee , Chaitanya Ahuja , Paul Pu Liang , Sanika Natu , Louis-Philippe Morency

Recent advancements in Large Video-Language Models (LVLMs) have led to promising results in multimodal video understanding. However, it remains unclear whether these models possess the cognitive capabilities required for high-level tasks,…

Computer Vision and Pattern Recognition · Computer Science 2025-07-02 Chenglin Li , Qianglong Chen , Zhi Li , Feng Tao , Yin Zhang

Memes have evolved as a prevalent medium for diverse communication, ranging from humour to propaganda. With the rising popularity of image-focused content, there is a growing need to explore its potential harm from different aspects.…

Computation and Language · Computer Science 2024-05-21 Siddhant Agarwal , Shivam Sharma , Preslav Nakov , Tanmoy Chakraborty

Multimodal/vision language models (VLMs) are increasingly being deployed in healthcare settings worldwide, necessitating robust benchmarks to ensure their safety, efficacy, and fairness. Multiple-choice question and answer (QA) datasets…

With the rapid development of large language models (LLMs) and their integration into large multimodal models (LMMs), there has been impressive progress in zero-shot completion of user-oriented vision-language tasks. However, a gap remains…

Computation and Language · Computer Science 2024-04-16 Fuxiao Liu , Xiaoyang Wang , Wenlin Yao , Jianshu Chen , Kaiqiang Song , Sangwoo Cho , Yaser Yacoob , Dong Yu

Embodied foundation models are increasingly performant in real-world domains such as robotics or autonomous driving. These models are often deployed in interactive or assistive settings, where it is important that these assistive models…

Robotics · Computer Science 2026-03-06 Pradyumna Tambwekar , Andrew Silva , Deepak Gopinath , Jonathan DeCastro , Xiongyi Cui , Guy Rosman

We propose a novel VQA dataset, BloomVQA, to facilitate comprehensive evaluation of large vision-language models on comprehension tasks. Unlike current benchmarks that often focus on fact-based memorization and simple reasoning tasks…

Computer Vision and Pattern Recognition · Computer Science 2024-06-11 Yunye Gong , Robik Shrestha , Jared Claypoole , Michael Cogswell , Arijit Ray , Christopher Kanan , Ajay Divakaran

Large Language Models (LLMs) have shown promise in the autonomous driving sector, particularly in generalization and interpretability. We introduce a unique object-level multimodal LLM architecture that merges vectorized numeric modalities…

With the rapid advancement of video generation models such as Sora, video quality assessment (VQA) is becoming increasingly crucial for selecting high-quality videos from large-scale datasets used in pre-training. Traditional VQA methods,…

Computer Vision and Pattern Recognition · Computer Science 2025-09-16 Yanyun Pu , Kehan Li , Zeyi Huang , Zhijie Zhong , Kaixiang Yang

Automatic analysis of teacher and student interactions could be very important to improve the quality of teaching and student engagement. However, despite some recent progress in utilizing multimodal data for teaching and learning…

Computers and Society · Computer Science 2022-12-07 Fangli Xu , Lingfei Wu , KP Thai , Carol Hsu , Wei Wang , Richard Tong

Human-designed visual manuals are crucial components in shape assembly activities. They provide step-by-step guidance on how we should move and connect different parts in a convenient and physically-realizable way. While there has been an…

Computer Vision and Pattern Recognition · Computer Science 2023-02-06 Ruocheng Wang , Yunzhi Zhang , Jiayuan Mao , Ran Zhang , Chin-Yi Cheng , Jiajun Wu

The growing volume of academic papers has made it increasingly difficult for researchers to efficiently extract key information. While large language models (LLMs) based agents are capable of automating question answering (QA) workflows for…

Computation and Language · Computer Science 2026-03-31 Tiancheng Huang , Ruisheng Cao , Yuxin Zhang , Zhangyi Kang , Zijian Wang , Chenrun Wang , Yijie Luo , Hang Zheng , Lirong Qian , Lu Chen , Kai Yu

Although numerous strategies have recently been proposed to enhance the autonomous interaction capabilities of multimodal agents in graphical user interface (GUI), their reliability remains limited when faced with complex or out-of-domain…

Computation and Language · Computer Science 2025-10-06 Pengzhou Cheng , Lingzhong Dong , Zeng Wu , Zongru Wu , Xiangru Tang , Chengwei Qin , Zhuosheng Zhang , Gongshen Liu

Within the multimodal field, large vision-language models (LVLMs) have made significant progress due to their strong perception and reasoning capabilities in the visual and language systems. However, LVLMs are still plagued by the two…

Computer Vision and Pattern Recognition · Computer Science 2024-06-14 Sirui Cheng , Siyu Zhang , Jiayi Wu , Muchen Lan

This paper presents an in-depth study of multimodal machine translation (MMT), examining the prevailing understanding that MMT systems exhibit decreased sensitivity to visual information when text inputs are complete. Instead, we attribute…

Computation and Language · Computer Science 2023-10-27 Yuxin Zuo , Bei Li , Chuanhao Lv , Tong Zheng , Tong Xiao , Jingbo Zhu

With the rapid growth in sensor data, effectively interpreting and interfacing with these data in a human-understandable way has become crucial. While existing research primarily focuses on learning classification models, fewer studies have…

Computation and Language · Computer Science 2025-03-04 Benjamin Reichman , Xiaofan Yu , Lanxiang Hu , Jack Truxal , Atishay Jain , Rushil Chandrupatla , Tajana Šimunić Rosing , Larry Heck

Real-time conversational assistants for procedural tasks often depend on video input, which can be computationally expensive and compromise user privacy. For the first time, we propose a real-time conversational assistant that provides…

Multimedia · Computer Science 2026-02-18 Rehana Mahfuz , Yinyi Guo , Erik Visser , Phanidhar Chinchili

Fusing sensors with complementary modalities is crucial for maintaining a stable and comprehensive understanding of abnormal driving scenes. However, Multimodal Large Language Models (MLLMs) are underexplored for leveraging multi-sensor…

Computer Vision and Pattern Recognition · Computer Science 2026-03-26 Mingzhe Tao , Ruiping Liu , Junwei Zheng , Yufan Chen , Kedi Ying , M. Saquib Sarfraz , Kailun Yang , Jiaming Zhang , Rainer Stiefelhagen

Action quality assessment (AQA) has become an emerging topic since it can be extensively applied in numerous scenarios. However, most existing methods and datasets focus on single-person short-sequence scenes, hindering the application of…

Computer Vision and Pattern Recognition · Computer Science 2024-04-09 Shiyi Zhang , Wenxun Dai , Sujia Wang , Xiangwei Shen , Jiwen Lu , Jie Zhou , Yansong Tang