English
Related papers

Related papers: Hierarchical Activity Recognition and Captioning f…

200 papers

There is a large variation in the activities that humans perform in their everyday lives. We consider modeling these composite human activities which comprises multiple basic level actions in a completely unsupervised setting. Our model…

Computer Vision and Pattern Recognition · Computer Science 2016-03-14 Chenxia Wu , Jiemi Zhang , Ozan Sener , Bart Selman , Silvio Savarese , Ashutosh Saxena

Daily Activity Recordings for Artificial Intelligence (DARai, pronounced "Dahr-ree") is a multimodal, hierarchically annotated dataset constructed to understand human activities in real-world settings. DARai consists of continuous scripted…

Computer Vision and Pattern Recognition · Computer Science 2026-03-30 Ghazal Kaviani , Yavuz Yarici , Seulgi Kim , Mohit Prabhushankar , Ghassan AlRegib , Mashhour Solh , Ameya Patil

Recent advances in Large Language Models (LLMs) and multimodal foundation models have significantly broadened their application in robotics and collaborative systems. However, effective multi-agent interaction necessitates robust…

Anticipating future actions is a highly challenging task due to the diversity and scale of potential future actions; yet, information from different modalities help narrow down plausible action choices. Each modality can provide diverse and…

Computer Vision and Pattern Recognition · Computer Science 2024-08-30 Apoorva Beedu , Harish Haresamudram , Karan Samel , Irfan Essa

This paper focuses on the temporal aspect for recognizing human activities in videos; an important visual cue that has long been undervalued. We revisit the conventional definition of activity and restrict it to Complex Action: a set of…

Computer Vision and Pattern Recognition · Computer Science 2019-04-30 Noureldien Hussein , Efstratios Gavves , Arnold W. M. Smeulders

In Psychology, actions are paramount for humans to identify sound events. In Machine Learning (ML), action recognition achieves high accuracy; however, it has not been asked whether identifying actions can benefit Sound Event Classification…

Sound · Computer Science 2021-08-09 Benjamin Elizalde , Radu Revutchi , Samarjit Das , Bhiksha Raj , Ian Lane , Laurie M. Heller

We present the Moments in Time Dataset, a large-scale human-annotated collection of one million short videos corresponding to dynamic events unfolding within three seconds. Modeling the spatial-audio-temporal dynamics even for actions…

Computer Vision and Pattern Recognition · Computer Science 2019-02-19 Mathew Monfort , Alex Andonian , Bolei Zhou , Kandan Ramakrishnan , Sarah Adel Bargal , Tom Yan , Lisa Brown , Quanfu Fan , Dan Gutfruend , Carl Vondrick , Aude Oliva

Recognizing Video events in long, complex videos with multiple sub-activities has received persistent attention recently. This task is more challenging than traditional action recognition with short, relatively homogeneous video clips. In…

Computer Vision and Pattern Recognition · Computer Science 2020-01-16 Yikang Li , Tianshu Yu , Baoxin Li

Automatic pronunciation assessment is a major component of a computer-assisted pronunciation training system. To provide in-depth feedback, scoring pronunciation at various levels of granularity such as phoneme, word, and utterance, with…

Computation and Language · Computer Science 2023-05-29 Heejin Do , Yunsu Kim , Gary Geunbae Lee

Human activities are particularly complex and variable, and this makes challenging for deep learning models to reason about them. However, we note that such variability does have an underlying structure, composed of a hierarchy of patterns…

Computer Vision and Pattern Recognition · Computer Science 2025-05-20 Simone Alberto Peirone , Francesca Pistilli , Giuseppe Averta

Abstractive conversation summarization has received much attention recently. However, these generated summaries often suffer from insufficient, redundant, or incorrect content, largely due to the unstructured and complex characteristics of…

Computation and Language · Computer Science 2021-04-20 Jiaao Chen , Diyi Yang

Dense video captioning is a fine-grained video understanding task that involves two sub-problems: localizing distinct events in a long video stream, and generating captions for the localized events. We propose the Joint Event Detection and…

Computer Vision and Pattern Recognition · Computer Science 2018-12-27 Huijuan Xu , Boyang Li , Vasili Ramanishka , Leonid Sigal , Kate Saenko

Current video generation models excel at creating short, realistic clips, but struggle with longer, multi-scene videos. We introduce \texttt{DreamFactory}, an LLM-based framework that tackles this challenge. \texttt{DreamFactory} leverages…

Artificial Intelligence · Computer Science 2024-08-22 Zhifei Xie , Daniel Tang , Dingwei Tan , Jacques Klein , Tegawend F. Bissyand , Saad Ezzini

Video caption refers to generating a descriptive sentence for a specific short video clip automatically, which has achieved remarkable success recently. However, most of the existing methods focus more on visual information while ignoring…

Computer Vision and Pattern Recognition · Computer Science 2017-12-12 Wangli Hao , Zhaoxiang Zhang , He Guan , Guibo Zhu

Video paragraph captioning (VPC) involves generating detailed narratives for long videos, utilizing supportive modalities such as speech and event boundaries. However, the existing models are constrained by the assumption of constant…

Computer Vision and Pattern Recognition · Computer Science 2024-03-29 Sishuo Chen , Lei Li , Shuhuai Ren , Rundong Gao , Yuanxin Liu , Xiaohan Bi , Xu Sun , Lu Hou

Automatic audio event recognition plays a pivotal role in making human robot interaction more closer and has a wide applicability in industrial automation, control and surveillance systems. Audio event is composed of intricate phonic…

Audio and Speech Processing · Electrical Eng. & Systems 2023-04-12 Tushar Sandhan , Sukanya Sonowal , Jin Young Choi

Multi-turn dialogues are characterized by their extended length and the presence of turn-taking conversations. Traditional language models often overlook the distinct features of these dialogues by treating them as regular text. In this…

Computation and Language · Computer Science 2024-02-01 Sangwoo Cho , Kaiqiang Song , Chao Zhao , Xiaoyang Wang , Dong Yu

We aim to automatically identify human action reasons in online videos. We focus on the widespread genre of lifestyle vlogs, in which people perform actions while verbally describing them. We introduce and make publicly available the WhyAct…

Computer Vision and Pattern Recognition · Computer Science 2021-09-10 Oana Ignat , Santiago Castro , Hanwen Miao , Weiji Li , Rada Mihalcea

We study the problem of directly deriving an initial human reenactment from a monocular video of a non-human character. Our goal is not to reconstruct the source character itself but to reinterpret its motion as a plausible and editable…

Computer Vision and Pattern Recognition · Computer Science 2026-05-20 Liuhan Chen , Lei Zhong , Jiewei Wang , Qin Shuai , Li Yuan , Leidong Fan , Qing Li , Kanglin Liu

Generating accurate descriptions of human actions in videos remains a challenging task for video captioning models. Existing approaches often struggle to capture fine-grained motion details, resulting in vague or semantically inconsistent…

Computer Vision and Pattern Recognition · Computer Science 2025-10-30 Guorui Song , Guocun Wang , Zhe Huang , Jing Lin , Xuefei Zhe , Jian Li , Haoqian Wang