English
Related papers

Related papers: Evoking Places from Spaces. The application of mul…

200 papers

We propose and study a novel cross-reality environment that seamlessly integrates a monoscopic 2D surface (an interactive screen with touch and pen input) with a stereoscopic 3D space (an augmented reality HMD) to jointly host spatial data…

Human-Computer Interaction · Computer Science 2024-09-25 Lixiang Zhao , Tobias Isenberg , Fuqi Xie , Hai-Ning Liang , Lingyun Yu

The multimedia communications with texts and images are popular on social media. However, limited studies concern how images are structured with texts to form coherent meanings in human cognition. To fill in the gap, we present a novel…

Multimedia · Computer Science 2023-02-28 Chunpu Xu , Hanzhuo Tan , Jing Li , Piji Li

In order for robots to operate effectively in homes and workplaces, they must be able to manipulate the articulated objects common within environments built for and by humans. Previous work learns kinematic models that prescribe this…

Robotics · Computer Science 2016-07-04 Zhengyang Wu , Mohit Bansal , Matthew R. Walter

Emergent narratives provide a unique and compelling approach to interactive storytelling through simulation, and have applications in games, narrative generation, and virtual agents. However the inherent complexity of simulation makes…

Artificial Intelligence · Computer Science 2020-04-24 Ben Kybartas , Clark Verbrugge , Jonathan Lessard

We propose Localized Narratives, a new form of multimodal image annotations connecting vision and language. We ask annotators to describe an image with their voice while simultaneously hovering their mouse over the region they are…

Computer Vision and Pattern Recognition · Computer Science 2020-07-21 Jordi Pont-Tuset , Jasper Uijlings , Soravit Changpinyo , Radu Soricut , Vittorio Ferrari

Story visualization aims to generate a sequence of images to narrate each sentence in a multi-sentence story, where the images should be realistic and keep global consistency across dynamic scenes and characters. Current works face the…

Computer Vision and Pattern Recognition · Computer Science 2022-11-15 Bowen Li , Thomas Lukasiewicz

The development of embodied agents that can communicate with humans in natural language has gained increasing interest over the last years, as it facilitates the diffusion of robotic platforms in human-populated environments. As a step…

Robotics · Computer Science 2024-04-16 Roberto Bigazzi , Marcella Cornia , Silvia Cascianelli , Lorenzo Baraldi , Rita Cucchiara

With the increased sophistication of AI techniques, the application of these systems has been expanding to ever newer fields. Increasingly, these systems are being used in modeling of human aesthetics and creativity, e.g. how humans create…

Computer Vision and Pattern Recognition · Computer Science 2019-09-17 Vanessa Utz , Steve DiPaola

Development of multimodal interactive systems is hindered by the lack of rich, multimodal (text, images) conversational data, which is needed in large quantities for LLMs. Previous approaches augment textual dialogues with retrieved images,…

Computation and Language · Computer Science 2024-10-04 Hossein Aboutalebi , Hwanjun Song , Yusheng Xie , Arshit Gupta , Justin Sun , Hang Su , Igor Shalyminov , Nikolaos Pappas , Siffi Singh , Saab Mansour

We present a universal framework to model contextualized sentence representations with visual awareness that is motivated to overcome the shortcomings of the multimodal parallel data with manual annotations. For each sentence, we first…

Computation and Language · Computer Science 2019-11-12 Zhuosheng Zhang , Rui Wang , Kehai Chen , Masao Utiyama , Eiichiro Sumita , Hai Zhao

Social concepts referring to non-physical objects--such as revolution, violence, or friendship--are powerful tools to describe, index, and query the content of visual data, including ever-growing collections of art images from the Cultural…

Computer Vision and Pattern Recognition · Computer Science 2026-04-20 Delfina Sol Martinez Pandiani , Valentina Presutti

In virtual reality (VR) educational scenarios, Pedagogical agents (PAs) enhance immersive learning through realistic appearances and interactive behaviors. However, most existing PAs rely on static speech and simple gestures. This…

Human-Computer Interaction · Computer Science 2026-03-11 Ninghao Wan , Jiarun Song , Fuzheng Yang

The human language can be expressed through multiple sources of information known as modalities, including tones of voice, facial gestures, and spoken language. Recent multimodal learning with strong performances on human-centric tasks such…

Computation and Language · Computer Science 2020-10-06 Yao-Hung Hubert Tsai , Martin Q. Ma , Muqiao Yang , Ruslan Salakhutdinov , Louis-Philippe Morency

Purpose: Surgical scene understanding is key to advancing computer-aided and intelligent surgical systems. Current approaches predominantly rely on visual data or end-to-end learning, which limits fine-grained contextual modeling. This work…

In this paper we present a novel interactive multimodal learning system, which facilitates search and exploration in large networks of social multimedia users. It allows the analyst to identify and select users of interest, and to find…

Information Retrieval · Computer Science 2019-05-08 Iva Gornishka , Stevan Rudinac , Marcel Worring

The main goal of this project is to research technical advances in order to enhance the possibility to develop narratives within immersive mediated environments. An important part of the research is concerned with the question of how a…

Human-Computer Interaction · Computer Science 2007-05-23 Joan Llobera

Many previous audio-visual voice-related works focus on speech, ignoring the singing voice in the growing number of musical video streams on the Internet. For processing diverse musical video data, voice activity detection is a necessary…

Sound · Computer Science 2021-06-23 Yuanbo Hou , Zhesong Yu , Xia Liang , Xingjian Du , Bilei Zhu , Zejun Ma , Dick Botteldooren

Multimodal deep learning has been used to predict clinical endpoints and diagnoses from clinical routine data. However, these models suffer from scaling issues: they have to learn pairwise interactions between each piece of information in…

This paper presents improved native unified multimodal models, \emph{i.e.,} Show-o2, that leverage autoregressive modeling and flow matching. Built upon a 3D causal variational autoencoder space, unified visual representations are…

Computer Vision and Pattern Recognition · Computer Science 2025-09-23 Jinheng Xie , Zhenheng Yang , Mike Zheng Shou

Humans rely on the synergy of their senses for most essential tasks. For tasks requiring object manipulation, we seamlessly and effectively exploit the complementarity of our senses of vision and touch. This paper draws inspiration from…

Robotics · Computer Science 2023-11-03 Carmelo Sferrazza , Younggyo Seo , Hao Liu , Youngwoon Lee , Pieter Abbeel