English
Related papers

Related papers: Identifying Visible Actions in Lifestyle Vlogs

200 papers

Video Action Detection (VAD) entails localizing and categorizing action instances within videos, which inherently consist of diverse information sources such as audio, visual cues, and surrounding scene contexts. Leveraging this multi-modal…

Computer Vision and Pattern Recognition · Computer Science 2025-02-04 Taein Son , Soo Won Seo , Jisong Kim , Seok Hwan Lee , Jun Won Choi

While most conversational AI systems focus on textual dialogue only, conditioning utterances on visual context (when it's available) can lead to more realistic conversations. Unfortunately, a major challenge for incorporating visual context…

Computer Vision and Pattern Recognition · Computer Science 2021-03-30 Paul Hongsuck Seo , Arsha Nagrani , Cordelia Schmid

This paper tackles the challenging task of evaluating socially situated conversational robots and presents a novel objective evaluation approach that relies on multimodal user behaviors. In this study, our main focus is on assessing the…

Computation and Language · Computer Science 2023-09-26 Koji Inoue , Divesh Lala , Keiko Ochi , Tatsuya Kawahara , Gabriel Skantze

Video captioning aims to describe video contents using natural language format that involves understanding and interpreting scenes, actions and events that occurs simultaneously on the view. Current approaches have mainly concentrated on…

Computer Vision and Pattern Recognition · Computer Science 2024-11-12 Antoine Hanna-Asaad , Decky Aspandi , Titus Zaharia

We present a novel approach for discovering human interactions in videos. Activity understanding techniques usually require a large number of labeled examples, which are not available in many practical cases. Here, we focus on recovering…

Computer Vision and Pattern Recognition · Computer Science 2015-02-16 Mehran Khodabandeh , Arash Vahdat , Guang-Tong Zhou , Hossein Hajimirsadeghi , Mehrsan Javan Roshtkhari , Greg Mori , Stephen Se

Given the enormous number of instructional videos available online, learning a diverse array of multi-step task models from videos is an appealing goal. We introduce a new pre-trained video model, VideoTaskformer, focused on representing…

Computer Vision and Pattern Recognition · Computer Science 2023-03-24 Medhini Narasimhan , Licheng Yu , Sean Bell , Ning Zhang , Trevor Darrell

The real-world is inherently multi-modal at its core. Our tools observe and take snapshots of it, in digital form, such as videos or sounds, however much of it is lost. Similarly for actions and information passing between humans, languages…

Computer Vision and Pattern Recognition · Computer Science 2026-05-26 Mihai-Cristian Pîrvu , Marius Leordeanu

Integrating higher level visual and linguistic interpretations is at the heart of human intelligence. As automatic visual category recognition in images is approaching human performance, the high level understanding in the dynamic…

Computer Vision and Pattern Recognition · Computer Science 2015-11-23 Anirudh Goyal , Marius Leordeanu

Social media has amplified the reach of financial influencers known as "finfluencers," who share stock recommendations on platforms like YouTube. Understanding their influence requires analyzing multimodal signals like tone, delivery style,…

Multimedia · Computer Science 2025-07-14 Michael Galarnyk , Veer Kejriwal , Agam Shah , Yash Bhardwaj , Nicholas Meyer , Anand Krishnan , Sudheer Chava

Large Language Models (LLMs) have allowed recent LLM-based approaches to achieve excellent performance on long-video understanding benchmarks. We investigate how extensive world knowledge and strong reasoning skills of underlying LLMs…

Computer Vision and Pattern Recognition · Computer Science 2025-06-12 Kanchana Ranasinghe , Xiang Li , Kumara Kahatapitiya , Michael S. Ryoo

In human vision objects and their parts can be visually recognized from purely spatial or purely temporal information but the mechanisms integrating space and time are poorly understood. Here we show that human visual recognition of objects…

Computer Vision and Pattern Recognition · Computer Science 2021-04-20 Guy Ben-Yosef , Gabriel Kreiman , Shimon Ullman

Human action recognition is an important problem in computer vision. It has a wide range of applications in surveillance, human-computer interaction, augmented reality, video indexing, and retrieval. The varying pattern of spatio-temporal…

Computer Vision and Pattern Recognition · Computer Science 2020-09-03 Yogesh S Rawat , Shruti Vyas

Multimodal machine translation involves drawing information from more than one modality, based on the assumption that the additional modalities will contain useful alternative views of the input data. The most prominent tasks in this area…

Computation and Language · Computer Science 2019-12-02 Umut Sulubacak , Ozan Caglayan , Stig-Arne Grönroos , Aku Rouhe , Desmond Elliott , Lucia Specia , Jörg Tiedemann

With the explosion of video content on the Internet, there is a need for research on methods for video analysis which take human cognition into account. One such cognitive measure is memorability, or the ability to recall visual content…

Computer Vision and Pattern Recognition · Computer Science 2017-08-29 Sumit Shekhar , Dhruv Singal , Harvineet Singh , Manav Kedia , Akhil Shetty

Realistic fake videos are a potential tool for spreading harmful misinformation given our increasing online presence and information intake. This paper presents a multimodal learning-based method for detection of real and fake videos. The…

Computer Vision and Pattern Recognition · Computer Science 2022-07-19 Kalin Stefanov , Bhawna Paliwal , Abhinav Dhall

Recognizing the activities causing distraction in real-world driving scenarios is critical for ensuring the safety and reliability of both drivers and pedestrians on the roadways. Conventional computer vision techniques are typically…

Computer Vision and Pattern Recognition · Computer Science 2024-03-22 Md Zahid Hasan , Jiajing Chen , Jiyang Wang , Mohammed Shaiqur Rahman , Ameya Joshi , Senem Velipasalar , Chinmay Hegde , Anuj Sharma , Soumik Sarkar

Gaze following estimates gaze targets of in-scene person by understanding human behavior and scene information. Existing methods usually analyze scene images for gaze following. However, compared with visual images, audio also provides…

Computer Vision and Pattern Recognition · Computer Science 2024-01-17 Yuqi Hou , Zhongqun Zhang , Nora Horanyi , Jaewon Moon , Yihua Cheng , Hyung Jin Chang

Video surveillance cameras generate most of recorded video, and there is far more recorded video than operators can watch. Much progress has recently been made using summarization of recorded video, but such techniques do not have much…

Computer Vision and Pattern Recognition · Computer Science 2017-01-05 Yedid Hoshen , Shmuel Peleg

Prior studies on Visual Sentiment Understanding (VSU) primarily rely on the explicit scene information (e.g., facial expression) to judge visual sentiments, which largely ignore implicit scene information (e.g., human action, objection…

Computer Vision and Pattern Recognition · Computer Science 2025-03-04 Jiamin Luo , Jingjing Wang , Junxiao Ma , Yujie Jin , Shoushan Li , Guodong Zhou

State-of-the-art temporal action detectors inefficiently search the entire video for specific actions. Despite the encouraging progress these methods achieve, it is crucial to design automated approaches that only explore parts of the video…

Computer Vision and Pattern Recognition · Computer Science 2018-07-30 Humam Alwassel , Fabian Caba Heilbron , Bernard Ghanem