中文
相关论文

相关论文: Multimodal Semantic Simulations of Linguistically …

200 篇论文

This paper addresses the generation of explanations with visual examples. Given an input sample, we build a system that not only classifies it to a specific category, but also outputs linguistic explanations and a set of visual examples…

计算机视觉与模式识别 · 计算机科学 2019-05-21 Atsushi Kanehira , Tatsuya Harada

We introduce the Scene Language, a visual scene representation that concisely and precisely describes the structure, semantics, and identity of visual scenes. It represents a scene with three key components: a program that specifies the…

计算机视觉与模式识别 · 计算机科学 2025-04-01 Yunzhi Zhang , Zizhang Li , Matt Zhou , Shangzhe Wu , Jiajun Wu

In this paper, we present an analysis of computationally generated mixed-modality definite referring expressions using combinations of gesture and linguistic descriptions. In doing so, we expose some striking formal semantic properties of…

计算与语言 · 计算机科学 2020-03-18 Nikhil Krishnaswamy , James Pustejovsky

We develop a system that formally represents spatial semantics concepts within natural language descriptions of spatial arrangements. The system builds on a model of spatial semantics representation according to which words in a sentence…

计算与语言 · 计算机科学 2021-11-30 Alexandros Haridis , Stella Rossikopoulou Pappa

In this paper we present a formal computational framework for modeling manipulation actions. The introduced formalism leads to semantics of manipulation action and has applications to both observing and understanding human manipulation…

机器人学 · 计算机科学 2015-12-07 Yezhou Yang , Yiannis Aloimonos , Cornelia Fermuller , Eren Erdal Aksoy

As robots begin to cohabit with humans in semi-structured environments, the need arises to understand instructions involving rich variability---for instance, learning to ground symbols in the physical world. Realistically, this task must…

人工智能 · 计算机科学 2017-06-02 Yordan Hristov , Svetlin Penkov , Alex Lascarides , Subramanian Ramamoorthy

In this paper, we argue that simulation platforms enable a novel type of embodied spatial reasoning, one facilitated by a formal model of object and event semantics that renders the continuous quantitative search space of an open-world,…

人工智能 · 计算机科学 2019-02-07 James Pustejovsky , Nikhil Krishnaswamy

In recent years, diffusion models have made remarkable strides in text-to-video generation, sparking a quest for enhanced control over video outputs to more accurately reflect user intentions. Traditional efforts predominantly focus on…

计算机视觉与模式识别 · 计算机科学 2024-07-18 Mingxiao Li , Bo Wan , Marie-Francine Moens , Tinne Tuytelaars

We explore using the Suggested Upper Merged Ontology (SUMO) to develop a semantic simulation. We provide two proof-of-concept demonstrations modeling transitions in a simulated gasoline engine using a general-purpose programming language.…

人工智能 · 计算机科学 2021-01-13 Robert B. Allen

How are we able to learn about complex current events just from short snippets of video? While natural language enables straightforward ways to represent under-specified, partially observable events, visual data does not facilitate…

Recent efforts to enable visual navigation using large language models have mainly focused on developing complex prompt systems. These systems incorporate instructions, observations, and history into massive text prompts, which are then…

计算机视觉与模式识别 · 计算机科学 2023-11-07 Yao-Hung Hubert Tsai , Vansh Dhar , Jialu Li , Bowen Zhang , Jian Zhang

High-dimensional distributed semantic spaces have proven useful and effective for aggregating and processing visual, auditory, and lexical information for many tasks related to human-generated data. Human language makes use of a large and…

计算与语言 · 计算机科学 2021-04-02 Jussi Karlgren , Pentti Kanerva

Inferring rigid-body physical states and properties from monocular videos is a fundamental step toward physics-based perception and simulation. Existing approaches assume specific underlying physical systems, object types, and camera poses,…

计算机视觉与模式识别 · 计算机科学 2026-05-21 Chia-Hsiang Kao , Cong Phuoc Huynh , Chien-Yi Wang , Noranart Vesdapunt , Stefan Stojanov , Bharath Hariharan , Oleksandr Obiednikov , Ning Zhou

With the goal of supporting scalable lexical semantic annotation, analysis, and theorizing, we conduct a comprehensive evaluation of different methods for generating event descriptions under both syntactic constraints -- e.g. desired clause…

计算与语言 · 计算机科学 2024-12-25 Angela Cao , Faye Holt , Jonas Chan , Stephanie Richter , Lelia Glass , Aaron Steven White

Computational simulations are a popular method for testing hypotheses about the emergence of communication. This kind of research is performed in a variety of traditions including language evolution, developmental psychology, cognitive…

人工智能 · 计算机科学 2023-03-09 Julian Zubek , Tomasz Korbak , Joanna Rączaszek-Leonardi

Language-guided scene-aware human motion generation has great significance for entertainment and robotics. In response to the limitations of existing datasets, we introduce LaserHuman, a pioneering dataset engineered to revolutionize…

计算机视觉与模式识别 · 计算机科学 2024-03-22 Peishan Cong , Ziyi Wang , Zhiyang Dou , Yiming Ren , Wei Yin , Kai Cheng , Yujing Sun , Xiaoxiao Long , Xinge Zhu , Yuexin Ma

Existing representations for human motion, such as MotionGPT, often operate as black-box latent vectors with limited interpretability and build on joint positions which can cause ambiguity. Inspired by the hierarchical structure of natural…

计算机视觉与模式识别 · 计算机科学 2026-03-17 Yao Zhang , Zhuchenyang Liu , Yu Xiao

Text-driven human motion generation is a multimodal task that synthesizes human motion sequences conditioned on natural language. It requires the model to satisfy textual descriptions under varying conditional inputs, while generating…

计算机视觉与模式识别 · 计算机科学 2024-10-01 Xingyu Chen

Open-vocabulary 3D scene understanding presents a significant challenge in computer vision, with wide-ranging applications in embodied agents and augmented reality systems. Existing methods adopt neurel rendering methods as 3D…

计算机视觉与模式识别 · 计算机科学 2024-08-26 Jun Guo , Xiaojian Ma , Yue Fan , Huaping Liu , Qing Li

Emotions play a central role in human communication, shaping trust, engagement, and social interaction. As artificial intelligence systems powered by large language models become increasingly integrated into everyday life, enabling them to…

音频与语音处理 · 电气工程与系统科学 2026-03-11 Soumya Dutta