中文
相关论文

相关论文: AMEGO: Active Memory from long EGOcentric videos

200 篇论文

Massive Open Online Courses (MOOCs) have become increasingly popular worldwide. However, learners primarily rely on watching videos, easily losing knowledge context and reducing learning effectiveness. We propose HyperMOOC, a novel approach…

人机交互 · 计算机科学 2026-02-02 Li Ye , Lei Wang , Lihong Cai , Ruiqi Yu , Yong Wang , Yigang Wang , Wei Chen , Zhiguang Zhou

In this work, we tackle the challenge of enhancing the realism and expressiveness in talking head video generation by focusing on the dynamic and nuanced relationship between audio cues and facial movements. We identify the limitations of…

计算机视觉与模式识别 · 计算机科学 2024-08-09 Linrui Tian , Qi Wang , Bang Zhang , Liefeng Bo

Egocentric videos can bring a lot of information about how humans perceive the world and interact with the environment, which can be beneficial for the analysis of human behaviour. The research in egocentric video analysis is developing…

计算机视觉与模式识别 · 计算机科学 2021-07-29 Ivan Rodin , Antonino Furnari , Dimitrios Mavroedis , Giovanni Maria Farinella

We present EgoACO, a deep neural architecture for video action recognition that learns to pool action-context-object descriptors from frame level features by leveraging the verb-noun structure of action labels in egocentric video datasets.…

计算机视觉与模式识别 · 计算机科学 2021-02-17 Swathikiran Sudhakaran , Sergio Escalera , Oswald Lanz

We present an approach for identifying picturesque highlights from large amounts of egocentric video data. Given a set of egocentric videos captured over the course of a vacation, our method analyzes the videos and looks for images that…

计算机视觉与模式识别 · 计算机科学 2016-01-19 Vinay Bettadapura , Daniel Castro , Irfan Essa

Video understanding has attracted much research attention especially since the recent availability of large-scale video benchmarks. In this paper, we address the problem of multi-label video classification. We first observe that there…

计算机视觉与模式识别 · 计算机科学 2017-11-07 Fang Yuan , Zhe Wang , Jie Lin , Luis Fernando D'Haro , Kim Jung Jae , Zeng Zeng , Vijay Chandrasekhar

The recent advancement of Vision Language Action (VLA) models has driven a critical demand for large scale egocentric datasets. However, existing datasets are often limited by short episode durations, typically spanning only a few minutes,…

计算机视觉与模式识别 · 计算机科学 2026-05-20 Senthil Palanisamy , Abhishek Anand , Satpal Singh Rathor , Pratyush Patnaik , Shubhanshu Khatana , Ekaksh Janweja

Recent multimodal large language models (MLLMs) have advanced video understanding, yet most still "think about videos" ie once a video is encoded, reasoning unfolds entirely in text, treating visual input as a static context. This passive…

计算机视觉与模式识别 · 计算机科学 2025-12-01 Hanoona Rasheed , Mohammed Zumri , Muhammad Maaz , Ming-Hsuan Yang , Fahad Shahbaz Khan , Salman Khan

The rapid development of Multimodal Large Language Models (MLLMs) has led to growing interest in egocentric video understanding, specifically the ability for MLLMs to recognize fine-grained hand-object interactions, track object state…

计算机视觉与模式识别 · 计算机科学 2026-05-20 Yang Dai , Dian Jiao , Tianwei Lin , Wenqiao Zhang

The goal of creating intelligent, human-centered wearable systems for continuous activity understanding faces a fundamental trade-off: Egocentric video-based models capture rich semantic information and have demonstrated strong performance…

计算机视觉与模式识别 · 计算机科学 2026-04-22 Baiyu Chen , Wilson Wongso , Zechen Li , Yonchanok Khaokaew , Hao Xue , Flora Salim

Recently, there has been a growing interest in wearable sensors which provides new research perspectives for 360 {\deg} video analysis. However, the lack of 360 {\deg} datasets in literature hinders the research in this field. To bridge…

计算机视觉与模式识别 · 计算机科学 2020-10-19 Keshav Bhandari , Mario A. DeLaGarza , Ziliang Zong , Hugo Latapie , Yan Yan

While affective computing has advanced considerably, multimodal emotion prediction in aging populations remains underexplored, largely due to the scarcity of dedicated datasets. Existing multimodal benchmarks predominantly target young,…

人机交互 · 计算机科学 2026-04-06 Hongbin Chen , Jie Li , Wei Wang , Siyang Song , Xiao Gu , Jianqing Li , Wentao Xiang

Long-horizon interactions between users and LLM-based assistants necessitate effective memory management, yet current approaches face challenges in training and evaluation of memory. Existing memory benchmarks rely on static, off-policy…

计算与语言 · 计算机科学 2026-03-03 Cheng Jiayang , Dongyu Ru , Lin Qiu , Yiyang Li , Xuezhi Cao , Yangqiu Song , Xunliang Cai

Multi-shot video generation requires maintaining a consistent appearance of recurring entities across shots while remaining faithful to shot-specific text prompts. Recent autoregressive methods reuse previously generated frames as memory.…

计算机视觉与模式识别 · 计算机科学 2026-05-25 Jente Vandersanden , Matheus Gadelha , Chun-Hao P. Huang , Hyeonho Jeong , Yulia Gryaditskaya

Visual queries 3D localization (VQ3D) is a task in the Ego4D Episodic Memory Benchmark. Given an egocentric video, the goal is to answer queries of the form "Where did I last see object X?", where the query object X is specified as a static…

计算机视觉与模式识别 · 计算机科学 2022-11-21 Jinjie Mai , Chen Zhao , Abdullah Hamdi , Silvio Giancola , Bernard Ghanem

Understanding human intentions (e.g., emotions) from videos has received considerable attention recently. Video streams generally constitute a blend of temporal data stemming from distinct modalities, including natural language, facial…

计算机视觉与模式识别 · 计算机科学 2024-10-01 Dingkang Yang , Mingcheng Li , Linhao Qu , Kun Yang , Peng Zhai , Song Wang , Lihua Zhang

As humans move around, performing their daily tasks, they are able to recall where they have positioned objects in their environment, even if these objects are currently out of their sight. In this paper, we aim to mimic this spatial…

计算机视觉与模式识别 · 计算机科学 2025-01-23 Chiara Plizzari , Shubham Goel , Toby Perrett , Jacob Chalk , Angjoo Kanazawa , Dima Damen

Anticipating actions before they occur is a core challenge in action understanding research. While conventional methods rely on extracting and aggregating temporal information from videos, as humans we can often predict upcoming actions by…

Understanding human tasks through video observations is an essential capability of intelligent agents. The challenges of such capability lie in the difficulty of generating a detailed understanding of situated actions, their effects on…

计算机视觉与模式识别 · 计算机科学 2022-10-11 Baoxiong Jia , Ting Lei , Song-Chun Zhu , Siyuan Huang

Multimodal large language models (MLLMs) are flourishing, but mainly focus on images with less attention than videos, especially in sub-fields such as prompt engineering, video chain-of-thought (CoT), and instruction tuning on videos.…

计算机视觉与模式识别 · 计算机科学 2024-07-09 Yan Wang , Yawen Zeng , Jingsheng Zheng , Xiaofen Xing , Jin Xu , Xiangmin Xu
‹ 上一页 1 8 9 10 下一页 ›