English
Related papers

Related papers: Video Dialog as Conversation about Objects Living …

200 papers

In this paper we present an approach and a benchmark for visual reasoning in robotics applications, in particular small object grasping and manipulation. The approach and benchmark are focused on inferring object properties from visual and…

Computer Vision and Pattern Recognition · Computer Science 2020-04-07 Michal Nazarczuk , Krystian Mikolajczyk

Video-grounded dialogue understanding is a challenging problem that requires machine to perceive, parse and reason over situated semantics extracted from weakly aligned video and dialogues. Most existing benchmarks treat both modalities the…

Computer Vision and Pattern Recognition · Computer Science 2023-05-31 Yuxuan Wang , Zilong Zheng , Xueliang Zhao , Jinpeng Li , Yueqian Wang , Dongyan Zhao

Dialog response ranking is used to rank response candidates by considering their relation to the dialog history. Although researchers have addressed this concept for open-domain dialogs, little attention has been focused on task-oriented…

Computation and Language · Computer Science 2018-11-29 Junki Ohmura , Maxine Eskenazi

Understanding and reasoning about dynamics governed by physical laws through visual observation, akin to human capabilities in the real world, poses significant challenges. Currently, object-centric dynamic simulation methods, which emulate…

Computer Vision and Pattern Recognition · Computer Science 2025-06-13 Jian Li , Wan Han , Ning Lin , Yu-Liang Zhan , Ruizhi Chengze , Haining Wang , Yi Zhang , Hongsheng Liu , Zidong Wang , Fan Yu , Hao Sun

Talking-head video editing aims to efficiently insert, delete, and substitute the word of a pre-recorded video through a text transcript editor. The key challenge for this task is obtaining an editing model that generates new talking-head…

Multimedia · Computer Science 2023-09-21 Songlin Yang , Wei Wang , Jun Ling , Bo Peng , Xu Tan , Jing Dong

We present the task of Spatio-Temporal Video Question Answering, which requires intelligent systems to simultaneously retrieve relevant moments and detect referenced visual concepts (people and objects) to answer natural language questions…

Computer Vision and Pattern Recognition · Computer Science 2020-05-13 Jie Lei , Licheng Yu , Tamara L. Berg , Mohit Bansal

Humans excel at multisensory perception and can often recognise object properties from the sound of their interactions. Inspired by this, we propose the novel task of Collision Sound Source Segmentation (CS3), where we aim to segment the…

Computer Vision and Pattern Recognition · Computer Science 2025-11-21 Kranti Kumar Parida , Omar Emara , Hazel Doughty , Dima Damen

Vision systems to see and reason about the compositional nature of visual scenes are fundamental to understanding our world. The complex relations between objects and their locations, ambiguities, and variations in the real-world…

Computer Vision and Pattern Recognition · Computer Science 2023-07-27 Muhammad Awais , Muzammal Naseer , Salman Khan , Rao Muhammad Anwer , Hisham Cholakkal , Mubarak Shah , Ming-Hsuan Yang , Fahad Shahbaz Khan

Dialogue state tracking (DST) aims to record user queries and goals during a conversational interaction achieved by maintaining a predefined set of slots and their corresponding values. Current approaches decide slot values opaquely, while…

Computation and Language · Computer Science 2024-03-12 Lin Xu , Ningxin Peng , Daquan Zhou , See-Kiong Ng , Jinlan Fu

Goal-oriented dialog systems enable users to complete specific goals like requesting information about a movie or booking a ticket. Typically the dialog system pipeline contains multiple ML models, including natural language understanding,…

We study the problem of dynamic visual reasoning on raw videos. This is a challenging problem; currently, state-of-the-art models often require dense supervision on physical object properties and events from simulation, which are…

Computer Vision and Pattern Recognition · Computer Science 2021-03-31 Zhenfang Chen , Jiayuan Mao , Jiajun Wu , Kwan-Yee Kenneth Wong , Joshua B. Tenenbaum , Chuang Gan

As the most critical components in a sentence, subject, predicate and object require special attention in the video captioning task. To implement this idea, we design a novel framework, named COllaborative three-Stream Transformers (COST),…

Computer Vision and Pattern Recognition · Computer Science 2023-09-19 Hao Wang , Libo Zhang , Heng Fan , Tiejian Luo

Socially competent robots should be equipped with the ability to perceive the world that surrounds them and communicate about it in a human-like manner. Representative skills that exhibit such ability include generating image descriptions…

Robotics · Computer Science 2021-02-01 Ting Han , Sina Zarrieß

Human conversation is a complex mechanism with subtle nuances. It is hence an ambitious goal to develop artificial intelligence agents that can participate fluently in a conversation. While we are still far from achieving this goal, recent…

Computer Vision and Pattern Recognition · Computer Science 2018-03-30 Unnat Jain , Svetlana Lazebnik , Alexander Schwing

Modelling and understanding time remains a challenge in contemporary video understanding models. With language emerging as a key driver towards powerful generalization, it is imperative for foundational video-language models to have a sense…

Computer Vision and Pattern Recognition · Computer Science 2023-03-28 Piyush Bagad , Makarand Tapaswi , Cees G. M. Snoek

With the recent advancements in AI, Intelligent Virtual Assistants (IVA) have become a ubiquitous part of every home. Going forward, we are witnessing a confluence of vision, speech and dialog system technologies that are enabling the IVAs…

Computation and Language · Computer Science 2018-12-21 Shachi H Kumar , Eda Okur , Saurav Sahay , Juan Jose Alvarado Leanos , Jonathan Huang , Lama Nachman

With robotics rapidly advancing, more effective human-robot interaction is increasingly needed to realize the full potential of robots for society. While spoken language must be part of the solution, our ability to provide spoken language…

Robotics · Computer Science 2024-03-25 Matthew Marge , Carol Espy-Wilson , Nigel Ward

We introduce ReXTime, a benchmark designed to rigorously test AI models' ability to perform temporal reasoning within video events. Specifically, ReXTime focuses on reasoning across time, i.e. human-like understanding when the question and…

Computer Vision and Pattern Recognition · Computer Science 2024-07-03 Jr-Jen Chen , Yu-Chien Liao , Hsi-Che Lin , Yu-Chu Yu , Yen-Chun Chen , Yu-Chiang Frank Wang

Image-to-video adaptation seeks to efficiently adapt image models for use in the video domain. Instead of finetuning the entire image backbone, many image-to-video adaptation paradigms use lightweight adapters for temporal modeling on top…

Computer Vision and Pattern Recognition · Computer Science 2024-07-10 Rui Qian , Shuangrui Ding , Dahua Lin

In this study, we introduce a low cost method for generating descriptions from images containing novel objects. Generally, constructing a model, which can explain images with novel objects, is costly because of the following: (1) collecting…

Computer Vision and Pattern Recognition · Computer Science 2020-03-09 Mikihiro Tanaka , Tatsuya Harada
‹ Prev 1 8 9 10 Next ›