中文
相关论文

相关论文: GEB+: A Benchmark for Generic Event Boundary Capti…

200 篇论文

While current video generation focuses on text or image conditions, practical applications like video editing and vlogging often need to seamlessly connect separate clips. In our work, we introduce Video Connecting, an innovative task that…

计算机视觉与模式识别 · 计算机科学 2026-01-28 Zhiyu Yin , Zhipeng Liu , Kehai Chen , Lemao Liu , Jin Liu , Hong-Dong Li , Yang Xiang , Min Zhang

There are substantial instructional videos on the Internet, which provide us tutorials for completing various tasks. Existing instructional video datasets only focus on specific steps at the video level, lacking experiential guidelines at…

计算机视觉与模式识别 · 计算机科学 2024-06-27 Jiafeng Liang , Shixin Jiang , Zekun Wang , Haojie Pan , Zerui Chen , Zheng Chu , Ming Liu , Ruiji Fu , Zhongyuan Wang , Bing Qin

We propose Visual News Captioner, an entity-aware model for the task of news image captioning. We also introduce Visual News, a large-scale benchmark consisting of more than one million news images along with associated news articles, image…

计算机视觉与模式识别 · 计算机科学 2021-09-15 Fuxiao Liu , Yinghan Wang , Tianlu Wang , Vicente Ordonez

The ability to perceive how objects change over time is a crucial ingredient in human intelligence. However, current benchmarks cannot faithfully reflect the temporal understanding abilities of video-language models (VidLMs) due to the…

计算机视觉与模式识别 · 计算机科学 2024-09-24 Shicheng Li , Lei Li , Shuhuai Ren , Yuanxin Liu , Yi Liu , Rundong Gao , Xu Sun , Lu Hou

We present EMBED (Egocentric Models Built with Exocentric Data), a method designed to transform exocentric video-language data for egocentric video representation learning. Large-scale exocentric data covers diverse activities with…

计算机视觉与模式识别 · 计算机科学 2024-08-08 Zi-Yi Dou , Xitong Yang , Tushar Nagarajan , Huiyu Wang , Jing Huang , Nanyun Peng , Kris Kitani , Fu-Jen Chu

Despite rapid advances in video generative models, robust metrics for evaluating visual and temporal correctness of complex human actions remain elusive. Critically, existing pure-vision encoders and Multimodal Large Language Models (MLLMs)…

计算机视觉与模式识别 · 计算机科学 2025-12-04 Xavier Thomas , Youngsun Lim , Ananya Srinivasan , Audrey Zheng , Deepti Ghadiyaram

Image captioning requires numerous annotated image-text pairs, resulting in substantial annotation costs. Recently, large models (e.g. diffusion models and large language models) have excelled in producing high-quality images and text. This…

计算机视觉与模式识别 · 计算机科学 2023-12-20 Feipeng Ma , Yizhou Zhou , Fengyun Rao , Yueyi Zhang , Xiaoyan Sun

Dense video understanding requires answering several questions such as who is doing what to whom, with what, how, why, and where. Recently, Video Situation Recognition (VidSitu) is framed as a task for structured prediction of multiple…

计算机视觉与模式识别 · 计算机科学 2022-10-21 Zeeshan Khan , C. V. Jawahar , Makarand Tapaswi

We generalize the notion of social biases from language embeddings to grounded vision and language embeddings. Biases are present in grounded embeddings, and indeed seem to be equally or more significant than for ungrounded embeddings. This…

计算与语言 · 计算机科学 2023-08-23 Candace Ross , Boris Katz , Andrei Barbu

We address the problem of video captioning by grounding language generation on object interactions in the video. Existing work mostly focuses on overall scene understanding with often limited or no emphasis on object interactions to address…

计算机视觉与模式识别 · 计算机科学 2017-11-20 Chih-Yao Ma , Asim Kadav , Iain Melvin , Zsolt Kira , Ghassan AlRegib , Hans Peter Graf

Complex Event Processing (CEP) is an event processing paradigm to perform real-time analytics over streaming data and match high-level event patterns. Presently, CEP is limited to process structured data stream. Video streams are…

计算机视觉与模式识别 · 计算机科学 2020-07-14 Piyush Yadav , Dhaval Salwala , Edward Curry

Video generation has many unique challenges beyond those of image generation. The temporal dimension introduces extensive possible variations across frames, over which consistency and continuity may be violated. In this study, we move…

计算机视觉与模式识别 · 计算机科学 2024-06-14 Weixi Feng , Jiachen Li , Michael Saxon , Tsu-jui Fu , Wenhu Chen , William Yang Wang

This work aims at generating captions for soccer videos using deep learning. In this context, this paper introduces a dataset, model, and triple-level evaluation. The dataset consists of 22k caption-clip pairs and three visual features…

计算机视觉与模式识别 · 计算机科学 2022-12-01 Ahmad Hammoudeh , Bastien Vanderplaetse , Stéphane Dupont

Text-level discourse parsing aims to unmask how two sentences in the text are related to each other. We propose the task of Visual Discourse Parsing, which requires understanding discourse relations among scenes in a video. Here we use the…

计算机视觉与模式识别 · 计算机科学 2022-01-25 Arjun R. Akula , Song-Chun Zhu

A generic video summary is an abridged version of a video that conveys the whole story and features the most important scenes. Yet the importance of scenes in a video is often subjective, and users should have the option of customizing the…

计算机视觉与模式识别 · 计算机科学 2021-12-09 Medhini Narasimhan , Anna Rohrbach , Trevor Darrell

We present PANDA, the first gigaPixel-level humAN-centric viDeo dAtaset, for large-scale, long-term, and multi-object visual analysis. The videos in PANDA were captured by a gigapixel camera and cover real-world scenes with both wide…

计算机视觉与模式识别 · 计算机科学 2020-03-11 Xueyang Wang , Xiya Zhang , Yinheng Zhu , Yuchen Guo , Xiaoyun Yuan , Liuyu Xiang , Zerun Wang , Guiguang Ding , David J Brady , Qionghai Dai , Lu Fang

There is growing interest in searching for information from large video corpora. Prior works have studied relevant tasks, such as text-based video retrieval, moment retrieval, video summarization, and video captioning in isolation, without…

计算机视觉与模式识别 · 计算机科学 2023-03-30 Abhay Zala , Jaemin Cho , Satwik Kottur , Xilun Chen , Barlas Oğuz , Yasher Mehdad , Mohit Bansal

Visual captioning benchmarks have become outdated with the emergence of modern multimodal large language models (MLLMs), as the brief ground-truth sentences and traditional metrics fail to assess detailed captions effectively. While recent…

计算机视觉与模式识别 · 计算机科学 2025-11-27 Zhihang Liu , Chen-Wei Xie , Bin Wen , Feiwu Yu , Jixuan Chen , Pandeng Li , Boqiang Zhang , Nianzu Yang , Yinglu Li , Zuan Gao , Yun Zheng , Hongtao Xie

Many real-world tasks require an agent to reason jointly over text and visual objects, (e.g., navigating in public spaces), which we refer to as context-sensitive text-rich visual reasoning. Specifically, these tasks require an…

计算机视觉与模式识别 · 计算机科学 2024-07-30 Rohan Wadhawan , Hritik Bansal , Kai-Wei Chang , Nanyun Peng

Detecting meaningful events in an untrimmed video is essential for dense video captioning. In this work, we propose a novel and simple model for event sequence generation and explore temporal relationships of the event sequence in the…

计算机视觉与模式识别 · 计算机科学 2020-06-16 Yuqing Song , Shizhe Chen , Yida Zhao , Qin Jin