English
Related papers

Related papers: PACS: A Dataset for Physical Audiovisual CommonSen…

200 papers

Humans are surrounded by audio signals that include both speech and non-speech sounds. The recognition and understanding of speech and non-speech audio events, along with a profound comprehension of the relationship between them, constitute…

Sound · Computer Science 2023-12-12 Yuan Gong , Alexander H. Liu , Hongyin Luo , Leonid Karlinsky , James Glass

Understanding the physical world - governed by laws of motion, spatial relations, and causality - poses a fundamental challenge for multimodal large language models (MLLMs). While recent advances such as OpenAI o3 and GPT-4o demonstrate…

Computer Vision and Pattern Recognition · Computer Science 2025-06-02 Zhuobai Dong , Junchao Yi , Ziyuan Zheng , Haochen Han , Xiangxi Zheng , Alex Jinpeng Wang , Fangming Liu , Linjie Li

The sciences of natural and artificial intelligence are fundamentally connected. Brain-inspired human-engineered AI are now the standard for predicting human brain responses during vision, and conversely, the brain continues to inspire…

Computer Vision and Pattern Recognition · Computer Science 2021-04-29 R. M. Cichy , K. Dwivedi , B. Lahner , A. Lascelles , P. Iamshchinina , M. Graumann , A. Andonian , N. A. R. Murty , K. Kay , G. Roig , A. Oliva

Recent advances in agentic AI have led to systems capable of autonomous task execution and language-based reasoning, yet their spatial reasoning abilities remain limited and underexplored, largely constrained to symbolic and sequential…

Artificial Intelligence · Computer Science 2025-09-12 Bui Duc Manh , Soumyaratna Debnath , Zetong Zhang , Shriram Damodaran , Arvind Kumar , Yueyi Zhang , Lu Mi , Erik Cambria , Lin Wang

A chief goal of artificial intelligence is to build machines that think like people. Yet it has been argued that deep neural network architectures fail to accomplish this. Researchers have asserted these models' limitations in the domains…

Machine Learning · Computer Science 2024-08-09 Luca M. Schulze Buschoff , Elif Akata , Matthias Bethge , Eric Schulz

Social intelligence is essential for understanding and reasoning about human expressions, intents and interactions. One representative benchmark for its study is Social Intelligence Queries (Social-IQ), a dataset of multiple-choice…

Computation and Language · Computer Science 2023-10-31 Xiao-Yu Guo , Yuan-Fang Li , Gholamreza Haffari

Understanding the physical world requires perceptual models grounded in physical laws rather than mere statistical correlations. However, existing multimodal learning frameworks, focused on vision and language, lack physical consistency and…

Artificial Intelligence · Computer Science 2025-11-26 Bo Pang , Chenxi Xu , Jierui Ren , Guoping Wang , Sheng Li

We introduce the new task of Acoustic Question Answering (AQA) to promote research in acoustic reasoning. The AQA task consists of analyzing an acoustic scene composed by a combination of elementary sounds and answering questions that…

Machine Learning · Computer Science 2019-03-01 Jerome Abdelnour , Giampiero Salvi , Jean Rouat

Language models have recently advanced into the realm of reasoning, yet it is through multimodal reasoning that we can fully unlock the potential to achieve more comprehensive, human-like cognitive capabilities. This survey provides a…

Computation and Language · Computer Science 2025-03-25 Zhiyu Lin , Yifei Gao , Xian Zhao , Yunfan Yang , Jitao Sang

Pre-trained language models (PTLMs) have achieved impressive performance on commonsense inference benchmarks, but their ability to employ commonsense to make robust inferences, which is crucial for effective communications with humans, is…

Computation and Language · Computer Science 2021-09-13 Pei Zhou , Rahul Khanna , Seyeon Lee , Bill Yuchen Lin , Daniel Ho , Jay Pujara , Xiang Ren

True intelligence hinges on the ability to uncover and leverage hidden causal relations. Despite significant progress in AI and computer vision (CV), there remains a lack of benchmarks for assessing models' abilities to infer latent…

Computer Vision and Pattern Recognition · Computer Science 2025-10-31 Disheng Liu , Yiran Qiao , Wuche Liu , Yiren Lu , Yunlai Zhou , Tuo Liang , Yu Yin , Jing Ma

Commonsense reasoning in multimodal contexts remains a foundational challenge in artificial intelligence. We introduce Multimodal UNcommonsense(MUN), a benchmark designed to evaluate models' ability to handle scenarios that deviate from…

Computer Vision and Pattern Recognition · Computer Science 2026-02-03 Yejin Son , Saejin Kim , Dongjun Min , Younjae Yu

Recent advancements in artificial intelligence have sparked interest in scientific assistants that could support researchers across the full spectrum of scientific workflows, from literature review to experimental design and data analysis.…

Existing dense or paragraph video captioning approaches rely on holistic representations of videos, possibly coupled with learned object/action representations, to condition hierarchical language decoders. However, they fundamentally lack…

Computer Vision and Pattern Recognition · Computer Science 2024-01-10 Shih-Han Chou , James J. Little , Leonid Sigal

Common sense has always been of interest in Artificial Intelligence, but has rarely taken center stage. Despite its mention in one of John McCarthy's earliest papers and years of work by dedicated researchers, arguably no AI system with a…

Artificial Intelligence · Computer Science 2022-02-08 Ronald J. Brachman , Hector J. Levesque

In this work, we introduce Contextual Analog Logic with Multimodality (CALM). CALM unites symbolic reasoning with neural generation, enabling systems to make context-sensitive decisions grounded in real-world multi-modal data. Background:…

Artificial Intelligence · Computer Science 2025-06-19 Maxwell J. Jacobson , Corey J. Maley , Yexiang Xue

Pre-trained on extensive text and image corpora, current Multi-Modal Large Language Models (MLLM) have shown strong capabilities in general visual reasoning tasks. However, their performance is still lacking in physical domains that require…

Artificial Intelligence · Computer Science 2025-07-04 Erle Zhu , Yadi Liu , Zhe Zhang , Xujun Li , Jin Zhou , Xinjie Yu , Minlie Huang , Hongning Wang

We propose a new task to benchmark human-in-scene understanding for embodied agents: Human-In-Scene Question Answering (HIS-QA). Given a human motion within a 3D scene, HIS-QA requires the agent to comprehend human states and behaviors,…

Computer Vision and Pattern Recognition · Computer Science 2025-07-17 Jiahe Zhao , Ruibing Hou , Zejie Tian , Hong Chang , Shiguang Shan

A comprehensive artificial intelligence system needs to not only perceive the environment with different `senses' (e.g., seeing and hearing) but also infer the world's conditional (or even causal) relations and corresponding uncertainty.…

Machine Learning · Statistics 2021-01-07 Hao Wang , Dit-Yan Yeung

Modern knowledge workplaces increasingly strain human episodic memory as individuals navigate fragmented attention, overlapping meetings, and multimodal information streams. Existing workplace tools provide partial support through…

Human-Computer Interaction · Computer Science 2026-03-03 Lawrence Obiuwevwi , Krzysztof J. Rechowicz , Vikas Ashok , Sachin Shetty , Sampath Jayarathna
‹ Prev 1 8 9 10 Next ›