English
Related papers

Related papers: D3G: Exploring Gaussian Prior for Temporal Sentenc…

200 papers

Dynamic scene rendering opens new avenues in autonomous driving by enabling closed-loop simulations with photorealistic data, which is crucial for validating end-to-end algorithms. However, the complex and highly dynamic nature of traffic…

Computer Vision and Pattern Recognition · Computer Science 2025-03-11 Rui Song , Chenwei Liang , Yan Xia , Walter Zimmer , Hu Cao , Holger Caesar , Andreas Festag , Alois Knoll

Spatio-temporal video grounding (STVG) aims to localize queried objects within dynamic video segments. Prevailing fully-trained approaches are notoriously data-hungry. However, gathering large-scale STVG data is exceptionally challenging:…

Computer Vision and Pattern Recognition · Computer Science 2026-04-15 Zanyi Wang , Fan Li , Dengyang Jiang , Liuzhuozheng Li , Yunhua Zhong , Guang Dai , Mengmeng Wang

3D Gaussian Splatting (3DGS) has emerged as a transformative method in the field of real-time novel synthesis. Based on 3DGS, recent advancements cope with large-scale scenes via spatial-based partition strategy to reduce video memory and…

Computer Vision and Pattern Recognition · Computer Science 2025-01-06 Tengfei Wang , Xin Wang , Yongmao Hou , Yiwei Xu , Wendi Zhang , Zongqian Zhan

Spatio-temporal grounding describes the task of localizing events in space and time, e.g., in video data, based on verbal descriptions only. Models for this task are usually trained with human-annotated sentences and bounding box…

Computer Vision and Pattern Recognition · Computer Science 2024-05-30 Brian Chen , Nina Shvetsova , Andrew Rouditchenko , Daniel Kondermann , Samuel Thomas , Shih-Fu Chang , Rogerio Feris , James Glass , Hilde Kuehne

3D Shape represented as point cloud has achieve advancements in multimodal pre-training to align image and language descriptions, which is curial to object identification, classification, and retrieval. However, the discrete representations…

Computer Vision and Pattern Recognition · Computer Science 2024-02-14 Haoyuan Li , Yanpeng Zhou , Yihan Zeng , Hang Xu , Xiaodan Liang

Video grounding aims to localize a moment from an untrimmed video for a given textual query. Existing approaches focus more on the alignment of visual and language stimuli with various likelihood-based matching or regression strategies,…

Computer Vision and Pattern Recognition · Computer Science 2021-07-08 Guoshun Nan , Rui Qiao , Yao Xiao , Jun Liu , Sicong Leng , Hao Zhang , Wei Lu

Due to the impressive zero-shot capabilities, pre-trained vision-language models (e.g., CLIP), have attracted widespread attention and adoption across various domains. Nonetheless, CLIP has been observed to be susceptible to adversarial…

Computer Vision and Pattern Recognition · Computer Science 2026-03-20 Lu Yu , Haiyang Zhang , Changsheng Xu

Dynamic scenes rendering is an intriguing yet challenging problem. Although current methods based on NeRF have achieved satisfactory performance, they still can not reach real-time levels. Recently, 3D Gaussian Splatting (3DGS) has garnered…

Computer Vision and Pattern Recognition · Computer Science 2024-11-07 Jiahao Lu , Jiacheng Deng , Ruijie Zhu , Yanzhe Liang , Wenfei Yang , Tianzhu Zhang , Xu Zhou

Video Question Answering (VideoQA) aims to answer natural language questions based on the information observed in videos. Despite the recent success of Large Multimodal Models (LMMs) in image-language understanding and reasoning, they deal…

Computer Vision and Pattern Recognition · Computer Science 2024-07-24 Haibo Wang , Chenghang Lai , Yixuan Sun , Weifeng Ge

Temporal grounding of text descriptions in videos is a central problem in vision-language learning and video understanding. Existing methods often prioritize accuracy over scalability -- they have been optimized for grounding only a few…

Computer Vision and Pattern Recognition · Computer Science 2024-04-08 Fangzhou Mu , Sicheng Mo , Yin Li

Semi-Supervised Video Paragraph Grounding (SSVPG) aims to localize multiple sentences in a paragraph from an untrimmed video with limited temporal annotations. Existing methods focus on teacher-student consistency learning and video-level…

Computer Vision and Pattern Recognition · Computer Science 2025-06-24 Yaokun Zhong , Siyu Jiang , Jian Zhu , Jian-Fang Hu

Video temporal grounding aims to identify video segments within untrimmed videos that are most relevant to a given natural language query. Existing video temporal localization models rely on specific datasets for training and have high data…

Computer Vision and Pattern Recognition · Computer Science 2024-08-30 Minghang Zheng , Xinhao Cai , Qingchao Chen , Yuxin Peng , Yang Liu

In this work we study Weakly Supervised Spatio-Temporal Video Grounding (WSTVG), a challenging task of localizing subjects spatio-temporally in videos using only textual queries and no bounding box supervision. Inspired by recent advances…

Computer Vision and Pattern Recognition · Computer Science 2025-04-08 Aaryan Garg , Akash Kumar , Yogesh S Rawat

Vision language action (VLA) models enable generalist robotic agents but often exhibit language ignorance, relying on visual shortcuts and remaining insensitive to instruction changes. We present Prospective Grounding and Alignment VLA…

Robotics · Computer Science 2026-04-14 Nastaran Darabi , Amit Ranjan Trivedi

Annotated datasets are critical for training neural networks for object detection, yet their manual creation is time- and labour-intensive, subjective to human error, and often limited in diversity. This challenge is particularly pronounced…

In this paper, we address a novel task, namely weakly-supervised spatio-temporally grounding natural sentence in video. Specifically, given a natural sentence and a video, we localize a spatio-temporal tube in the video that semantically…

Computer Vision and Pattern Recognition · Computer Science 2019-06-07 Zhenfang Chen , Lin Ma , Wenhan Luo , Kwan-Yee K. Wong

While visual-language models have profoundly linked features between texts and images, the incorporation of 3D modality data, such as point clouds and 3D Gaussians, further enables pretraining for 3D-related tasks, e.g., cross-modal…

Computer Vision and Pattern Recognition · Computer Science 2026-01-28 Jiarun Liu , Qifeng Chen , Yiru Zhao , Minghua Liu , Baorui Ma , Sheng Yang

Gradient temporal difference (Gradient TD) algorithms are a popular class of stochastic approximation (SA) algorithms used for policy evaluation in reinforcement learning. Here, we consider Gradient TD algorithms with an additional heavy…

Machine Learning · Computer Science 2021-11-23 Rohan Deb , Shalabh Bhatnagar

Inspired by the activity-silent and persistent activity mechanisms in human visual perception biology, we design a Unified Static and Dynamic Network (UniSDNet), to learn the semantic association between the video and text/audio queries in…

Computer Vision and Pattern Recognition · Computer Science 2025-04-14 Jingjing Hu , Dan Guo , Kun Li , Zhan Si , Xun Yang , Xiaojun Chang , Meng Wang

Due to the lack of temporal annotation, current Weakly-supervised Temporal Action Localization (WTAL) methods are generally stuck into over-complete or incomplete localization. In this paper, we aim to leverage the text information to boost…

Computer Vision and Pattern Recognition · Computer Science 2023-05-02 Guozhang Li , De Cheng , Xinpeng Ding , Nannan Wang , Xiaoyu Wang , Xinbo Gao
‹ Prev 1 4 5 6 7 8 10 Next ›