中文
相关论文

相关论文: A Simple Transformer-Based Model for Ego4D Natural…

200 篇论文

Large language Models (LLMs) are usually used to answer questions, but many high-stakes applications (e.g., tutoring, clinical support) require the complementary skill of asking questions: detecting missing information, requesting…

人工智能 · 计算机科学 2026-01-07 Rajeev Bhatt Ambati , Tianyi Niu , Aashu Singh , Shlok Mishra , Snigdha Chaturvedi , Shashank Srivastava

Video Referring Expression Comprehension (REC) aims to localize a target object in videos based on the queried natural language. Recent improvements in video REC have been made using Transformer-based methods with learnable queries.…

计算机视觉与模式识别 · 计算机科学 2023-10-26 Ji Jiang , Meng Cao , Tengtao Song , Long Chen , Yi Wang , Yuexian Zou

With the recent advances in video and 3D understanding, novel 4D spatio-temporal methods fusing both concepts have emerged. Towards this direction, the Ego4D Episodic Memory Benchmark proposed a task for Visual Queries with 3D Localization…

计算机视觉与模式识别 · 计算机科学 2023-08-29 Jinjie Mai , Abdullah Hamdi , Silvio Giancola , Chen Zhao , Bernard Ghanem

In episodic memory with natural language queries (EM-NLQ), a user may ask a question (e.g., "Where did I place the mug?") that requires searching a long egocentric video, captured from the user's perspective, to find the moment that answers…

计算机视觉与模式识别 · 计算机科学 2026-04-29 Nikesh Subedi , Loris Bazzani , Ziad Al-Halah

Recent years witnessed an increase in the amount of research on the task of Question Difficulty Estimation from Text QDET with Natural Language Processing (NLP) techniques, with the goal of targeting the limitations of traditional…

计算与语言 · 计算机科学 2023-05-18 Luca Benedetto

In egocentric videos, actions occur in quick succession. We capitalise on the action's temporal context and propose a method that learns to attend to surrounding actions in order to improve recognition performance. To incorporate the…

计算机视觉与模式识别 · 计算机科学 2021-11-02 Evangelos Kazakos , Jaesung Huh , Arsha Nagrani , Andrew Zisserman , Dima Damen

In this paper, we present our contribution to ABAW facial expression challenge. We report the proposed system and the official challenge results adhering to the challenge protocol. Using end-to-end deep learning and benefiting from transfer…

机器学习 · 计算机科学 2020-11-03 Denis Dresvyanskiy , Elena Ryumina , Heysem Kaya , Maxim Markitantov , Alexey Karpov , Wolfgang Minker

In this report, we describe our submission to the Ego4D AudioVisual (AV) Speech Transcription Challenge 2022. Our pipeline is based on AVATAR, a state of the art encoder-decoder model for AV-ASR that performs early fusion of spectrograms…

计算机视觉与模式识别 · 计算机科学 2022-11-21 Paul Hongsuck Seo , Arsha Nagrani , Cordelia Schmid

Video understanding typically requires fine-tuning the large backbone when adapting to new domains. In this paper, we leverage the egocentric video foundation models (Ego-VFMs) based on video-language pre-training and propose a…

计算机视觉与模式识别 · 计算机科学 2025-02-28 Tz-Ying Wu , Kyle Min , Subarna Tripathi , Nuno Vasconcelos

Automatic emotion recognition is a challenging task. In this paper, we present our effort for the audio-video based sub-challenge of the Emotion Recognition in the Wild (EmotiW) 2018 challenge, which requires participants to assign a single…

计算机视觉与模式识别 · 计算机科学 2018-09-18 Zheng Lian , Ya Li , Jianhua Tao , Jian Huang

Understanding videos to localize moments with natural language often requires large expensive annotated video regions paired with language queries. To eliminate the annotation costs, we make a first attempt to train a natural language video…

计算与语言 · 计算机科学 2021-10-04 Jinwoo Nam , Daechul Ahn , Dongyeop Kang , Seong Jong Ha , Jonghyun Choi

We investigate whether off-the-shelf Multimodal Large Language Models (MLLMs) can tackle Online Episodic-Memory Video Question Answering (OEM-VQA) without additional training. Our pipeline converts a streaming egocentric video into a…

计算机视觉与模式识别 · 计算机科学 2025-06-23 Giuseppe Lando , Rosario Forte , Giovanni Maria Farinella , Antonino Furnari

Transformer based architectures have shown notable results on many down streaming tasks including question answering. The availability of data, on the other hand, impedes obtaining legitimate performance for low-resource languages. In this…

计算与语言 · 计算机科学 2024-09-04 Hariom A. Pandya , Bhavik Ardeshna , Brijesh S. Bhatt

Interaction and navigation defined by natural language instructions in dynamic environments pose significant challenges for neural agents. This paper focuses on addressing two challenges: handling long sequence of subtasks, and…

计算机视觉与模式识别 · 计算机科学 2021-08-26 Alexander Pashevich , Cordelia Schmid , Chen Sun

Transformers have achieved remarkable performance in a myriad of fields including natural language processing and computer vision. However, when it comes to the graph mining area, where graph neural network (GNN) has been the dominant…

机器学习 · 计算机科学 2021-10-26 Jianan Zhao , Chaozhuo Li , Qianlong Wen , Yiqi Wang , Yuming Liu , Hao Sun , Xing Xie , Yanfang Ye

Text-to-Image generation in the general domain has long been an open problem, which requires both a powerful generative model and cross-modal understanding. We propose CogView, a 4-billion-parameter Transformer with VQ-VAE tokenizer to…

计算机视觉与模式识别 · 计算机科学 2021-11-08 Ming Ding , Zhuoyi Yang , Wenyi Hong , Wendi Zheng , Chang Zhou , Da Yin , Junyang Lin , Xu Zou , Zhou Shao , Hongxia Yang , Jie Tang

Recently there has been a rising interest in training agents, embodied in virtual environments, to perform language-directed tasks by deep reinforcement learning. In this paper, we propose a simple but effective neural language grounding…

人工智能 · 计算机科学 2018-09-06 Haonan Yu , Xiaochen Lian , Haichao Zhang , Wei Xu

Deep neural networks facilitate video question answering (VideoQA), but the real-world applications on video streams such as CCTV and live cast place higher demands on the solver. To address the challenges of VideoQA on long videos of…

多媒体 · 计算机科学 2023-03-08 Weikai Kong , Shuhong Ye , Chenglin Yao , Jianfeng Ren

The COVID-19 pandemic and the internet's availability have recently boosted online learning. However, monitoring engagement in online learning is a difficult task for teachers. In this context, timely automatic student engagement…

计算机视觉与模式识别 · 计算机科学 2025-02-18 Sandeep Mandia , Kuldeep Singh , Rajendra Mitharwal , Faisel Mushtaq , Dimpal Janu

This paper proposes the MT-DQN model, which integrates a Transformer, Temporal Graph Neural Network (TGNN), and Deep Q-Network (DQN) to address the challenges of predicting user behavior and optimizing recommendation strategies in…

机器学习 · 计算机科学 2025-09-17 Jinmeiyang Wang , Jing Dong , Li Zhou