中文
相关论文

相关论文: ProMQA-Assembly: Multimodal Procedural QA Dataset …

200 篇论文

Shape assembly is a ubiquitous task in daily life, integral for constructing complex 3D structures like IKEA furniture. While significant progress has been made in developing autonomous agents for shape assembly, existing datasets have not…

计算机视觉与模式识别 · 计算机科学 2024-11-19 Yunong Liu , Cristobal Eyzaguirre , Manling Li , Shubh Khanna , Juan Carlos Niebles , Vineeth Ravi , Saumitra Mishra , Weiyu Liu , Jiajun Wu

It is very challenging to curate a dataset for language-specific knowledge and common sense in order to evaluate natural language understanding capabilities of language models. Due to the limitation in the availability of annotators, most…

计算与语言 · 计算机科学 2024-06-07 Yusuke Sakai , Hidetaka Kamigaito , Taro Watanabe

Machine reading is a fundamental task for testing the capability of natural language understanding, which is closely related to human cognition in many aspects. With the rising of deep learning techniques, algorithmic models rival human…

计算与语言 · 计算机科学 2020-07-17 Jian Liu , Leyang Cui , Hanmeng Liu , Dandan Huang , Yile Wang , Yue Zhang

Constructing scientific multimodal document reasoning datasets for foundation model training involves an inherent trade-off among scale, faithfulness, and realism. To address this challenge, we introduce the synthesize-and-reground…

计算与语言 · 计算机科学 2026-04-30 Ziyu Chen , Yilun Zhao , Chengye Wang , Rilyn Han , Manasi Patwardhan , Arman Cohan

Question-answering (QA) is a natural approach for humans to understand a piece of music audio. However, for machines, accessing a large-scale dataset covering diverse aspects of music is crucial, yet challenging, due to the scarcity of…

声音 · 计算机科学 2025-08-28 Zhihao Ouyang , Ju-Chiang Wang , Daiyu Zhang , Bin Chen , Shangjie Li , Quan Lin

Training multimodal agents via reinforcement learning for knowledge-intensive visual reasoning is fundamentally hindered by the extreme sparsity of outcome-based supervision and the unpredictability of live web environments. To resolve…

计算机视觉与模式识别 · 计算机科学 2026-04-23 Wentao Yan , Shengqin Wang , Huichi Zhou , Yihang Chen , Kun Shao , Yuan Xie , Zhizhong Zhang

The evolution of digital manufacturing requires intelligent Question Answering (QA) systems that can seamlessly integrate and analyze complex multi-modal data, such as text, images, formulas, and tables. Conventional Retrieval Augmented…

计算工程、金融与科学 · 计算机科学 2026-01-27 Yunqing Li , Zihan Dong , Farhad Ameri , Jianbang Zhang

In this paper, we introduce the MLM (Multiple Languages and Modalities) dataset - a new resource to train and evaluate multitask systems on samples in multiple modalities and three languages. The generation process and inclusion of semantic…

机器学习 · 计算机科学 2020-10-27 Jason Armitage , Endri Kacupaj , Golsa Tahmasebzadeh , Swati , Maria Maleshkova , Ralph Ewerth , Jens Lehmann

Biomedical data is inherently multimodal, consisting of electronic health records, medical imaging, digital pathology, genome sequencing, wearable sensors, and more. The application of artificial intelligence tools to these multifaceted…

机器学习 · 计算机科学 2024-08-26 Shentong Mo , Paul Pu Liang

We present VQA-MHUG - a novel 49-participant dataset of multimodal human gaze on both images and questions during visual question answering (VQA) collected using a high-speed eye tracker. We use our dataset to analyze the similarity between…

计算机视觉与模式识别 · 计算机科学 2026-03-04 Ekta Sood , Fabian Kögel , Florian Strohm , Prajit Dhar , Andreas Bulling

Next generation task-oriented dialog systems need to understand conversational contexts with their perceived surroundings, to effectively help users in the real-world multimodal environment. Existing task-oriented dialog datasets aimed…

计算与语言 · 计算机科学 2021-10-22 Satwik Kottur , Seungwhan Moon , Alborz Geramifard , Babak Damavandi

When answering complex questions, people can seamlessly combine information from visual, textual and tabular sources. While interest in models that reason over multiple pieces of evidence has surged in recent years, there has been…

Wearable systems can recognize activities from IMU data but often fail to explain their underlying causes or contextual significance. To address this limitation, we introduce two large-scale resources: SensorCap, comprising 35,960…

计算与语言 · 计算机科学 2025-09-23 Sheikh Asif Imran , Mohammad Nur Hossain Khan , Subrata Biswas , Bashima Islam

Multimodal Large Language Models (MLLMs) are evolving from passive observers into active agents, solving problems through Visual Expansion (invoking visual tools) and Knowledge Expansion (open-web search). However, existing evaluations fall…

Spoken question answering (SQA) systems are critical for digital assistants and other real-world use cases, but evaluating their performance is a challenge due to the importance of human-spoken questions. This study presents a new…

In recent years, the advancement of AI technologies has accelerated the development of smart factories. In particular, the automatic monitoring of product assembly progress is crucial for improving operational efficiency, minimizing the…

计算机视觉与模式识别 · 计算机科学 2026-01-05 Kazuma Miura , Sarthak Pathak , Kazunori Umeda

We present TriviaQA, a challenging reading comprehension dataset containing over 650K question-answer-evidence triples. TriviaQA includes 95K question-answer pairs authored by trivia enthusiasts and independently gathered evidence…

计算与语言 · 计算机科学 2017-05-16 Mandar Joshi , Eunsol Choi , Daniel S. Weld , Luke Zettlemoyer

The ongoing trend towards Industry 4.0 has revolutionised ordinary workplaces, profoundly changing the role played by humans in the production chain. Research on ergonomics in industrial settings mainly focuses on reducing the operator's…

机器人学 · 计算机科学 2022-07-11 Marta Lagomarsino , Marta Lorenzini , Elena De Momi , Arash Ajoudani

Walking assistance in extreme or complex environments remains a significant challenge for people with blindness or low vision (BLV), largely due to the lack of a holistic scene understanding. Motivated by the real-world needs of the BLV…

计算机视觉与模式识别 · 计算机科学 2025-10-24 Kedi Ying , Ruiping Liu , Chongyan Chen , Mingzhe Tao , Hao Shi , Kailun Yang , Jiaming Zhang , Rainer Stiefelhagen

This work presents the Industrial Hand Action Dataset V1, an industrial assembly dataset consisting of 12 classes with 459,180 images in the basic version and 2,295,900 images after spatial augmentation. Compared to other freely available…

计算机视觉与模式识别 · 计算机科学 2023-03-08 Fabian Sturm , Elke Hergenroether , Julian Reinhardt , Petar Smilevski Vojnovikj , Melanie Siegel