中文
相关论文

相关论文: ShotVL: Human-Centric Highlight Frame Retrieval vi…

200 篇论文

Precise action localization in untrimmed video is vital for fields such as professional sports and minimally invasive surgery, where the delineation of particular motions in recordings can dramatically enhance analysis. But in many cases,…

计算机视觉与模式识别 · 计算机科学 2025-03-26 Josiah Aklilu , Xiaohan Wang , Serena Yeung-Levy

Large scale Vision-Language (VL) models have shown tremendous success in aligning representations between visual and text modalities. This enables remarkable progress in zero-shot recognition, image generation & editing, and many other…

计算机视觉与模式识别 · 计算机科学 2023-07-25 Wei Lin , Leonid Karlinsky , Nina Shvetsova , Horst Possegger , Mateusz Kozinski , Rameswar Panda , Rogerio Feris , Hilde Kuehne , Horst Bischof

In this work, we introduce a novel task - Humancentric Spatio-Temporal Video Grounding (HC-STVG). Unlike the existing referring expression tasks in images or videos, by focusing on humans, HC-STVG aims to localize a spatiotemporal tube of…

计算机视觉与模式识别 · 计算机科学 2021-06-03 Zongheng Tang , Yue Liao , Si Liu , Guanbin Li , Xiaojie Jin , Hongxu Jiang , Qian Yu , Dong Xu

We propose a new "Unbiased through Textual Description (UTD)" video benchmark based on unbiased subsets of existing video classification and retrieval datasets to enable a more robust assessment of video understanding capabilities. Namely,…

计算机视觉与模式识别 · 计算机科学 2025-03-25 Nina Shvetsova , Arsha Nagrani , Bernt Schiele , Hilde Kuehne , Christian Rupprecht

Reference-based video object segmentation is an emerging topic which aims to segment the corresponding target object in each video frame referred by a given reference, such as a language expression or a photo mask. However, language…

计算机视觉与模式识别 · 计算机科学 2023-11-14 Ruolin Yang , Da Li , Conghui Hu , Timothy Hospedales , Honggang Zhang , Yi-Zhe Song

Multimodal Large Language Models (MLLMs) have demonstrated significant advances in visual understanding tasks. However, their capacity to comprehend human-centric scenes has rarely been explored, primarily due to the absence of…

Despite great recent advances in visual tracking, its further development, including both algorithm design and evaluation, is limited due to lack of dedicated large-scale benchmarks. To address this problem, we present LaSOT, a high-quality…

计算机视觉与模式识别 · 计算机科学 2020-09-15 Heng Fan , Hexin Bai , Liting Lin , Fan Yang , Peng Chu , Ge Deng , Sijia Yu , Harshit , Mingzhen Huang , Juehuan Liu , Yong Xu , Chunyuan Liao , Lin Yuan , Haibin Ling

Vision-Language Models (VLMs) have demonstrated remarkable capabilities in general video understanding, yet they often struggle with the fine-grained comprehension crucial for real-world applications requiring nuanced interpretation of…

计算机视觉与模式识别 · 计算机科学 2026-05-26 Gueter Josmy Faure , Min-Hung Chen , Jia-Fong Yeh , Hung-Ting Su , Winston H. Hsu

Humans excel at visual social inference, the ability to infer hidden elements of a scene from subtle behavioral cues such as other people's gaze, pose, and orientation. This ability drives everyday social reasoning in humans and is critical…

计算机视觉与模式识别 · 计算机科学 2025-11-20 Neha Balamurugan , Sarah Wu , Adam Chun , Gabe Gaw , Cristobal Eyzaguirre , Tobias Gerstenberg

Robotic search of people in human-centered environments, including healthcare settings, is challenging as autonomous robots need to locate people without complete or any prior knowledge of their schedules, plans or locations. Furthermore,…

机器人学 · 计算机科学 2024-12-03 Angus Fung , Aaron Hao Tan , Haitong Wang , Beno Benhabib , Goldie Nejat

Video summarization has become an increasingly important task in the field of computer vision due to the vast amount of video content available on the internet. In this project, we propose a new method for natural language query based joint…

计算机视觉与模式识别 · 计算机科学 2023-05-10 Richard Luo , Austin Peng , Heidi Yap , Koby Beard

Following human instructions to explore and search for a specified target in an unfamiliar environment is a crucial skill for mobile service robots. Most of the previous works on object goal navigation have typically focused on a single…

机器人学 · 计算机科学 2024-11-19 Bangguo Yu , Yuzhen Liu , Lei Han , Hamidreza Kasaei , Tingguang Li , Ming Cao

Long-form multimodal video understanding requires integrating vision, speech, and ambient audio with coherent long-range reasoning. Existing benchmarks emphasize either temporal length or multimodal richness, but rarely both and while some…

In this paper, we present LOC-ZSON, a novel Language-driven Object-Centric image representation for object navigation task within complex scenes. We propose an object-centric image representation and corresponding losses for visual-language…

计算机视觉与模式识别 · 计算机科学 2024-05-10 Tianrui Guan , Yurou Yang , Harry Cheng , Muyuan Lin , Richard Kim , Rajasimman Madhivanan , Arnie Sen , Dinesh Manocha

Key-value relations are prevalent in Visually-Rich Documents (VRDs), often depicted in distinct spatial regions accompanied by specific color and font styles. These non-textual cues serve as important indicators that greatly enhance human…

计算机视觉与模式识别 · 计算机科学 2024-03-26 Hao Wang , Tang Li , Chenhui Chu , Nengjun Zhu , Rui Wang , Pinpin Zhu

Video quality is a primary concern for video service providers. In recent years, the techniques of video quality assessment (VQA) based on deep convolutional neural networks (CNNs) have been developed rapidly. Although existing works…

图像与视频处理 · 电气工程与系统科学 2022-10-11 Ao-Xiang Zhang , Yuan-Gen Wang , Weixuan Tang , Leida Li , Sam Kwong

Shot language understanding (SLU) is crucial for cinematic analysis but remains challenging due to its diverse cinematographic dimensions and subjective expert judgment. While vision-language models (VLMs) have shown strong ability in…

机器学习 · 计算机科学 2026-03-20 Haoxin Liu , Harshavardhan Kamarthi , Zhiyuan Zhao , Hongjie Chen , B. Aditya Prakash

Inspired by recent advances in neural machine translation, that jointly align and translate using encoder-decoder networks equipped with attention, we propose an attentionbased LSTM model for human activity recognition. Our model jointly…

计算机视觉与模式识别 · 计算机科学 2017-09-01 Atousa Torabi , Leonid Sigal

Vision-language models (VLMs) trained on internet-scale data achieve remarkable zero-shot detection performance on common objects like car, truck, and pedestrian. However, state-of-the-art models still struggle to generalize to…

计算机视觉与模式识别 · 计算机科学 2025-10-24 Peter Robicheaux , Matvei Popov , Anish Madan , Isaac Robinson , Joseph Nelson , Deva Ramanan , Neehar Peri

Template-based 3D object tracking still lacks a high-precision benchmark of real scenes due to the difficulty of annotating the accurate 3D poses of real moving video objects without using markers. In this paper, we present a multi-view…

计算机视觉与模式识别 · 计算机科学 2022-03-28 Jiachen Li , Bin Wang , Shiqiang Zhu , Xin Cao , Fan Zhong , Wenxuan Chen , Te Li , Jason Gu , Xueying Qin